python实现读取pdf格式文档
一、 准备工作
安装对应的库
pip install pdfminer3k
pip install pdfminer.six
二、部分变量的含义
PDFDocument(pdf文档对象)
PDFPageInterpreter(解释器)
PDFParser(pdf文档分析器)
PDFResourceManager(资源管理器)
PDFPageAggregator(聚合器)
LAParams(参数分析器)
三、PDFMiner类之间的关系

PDFMiner的相关文档(点击跳转)
四、代码实现
def changePdfToText(filePath):
"""
解析pdf 文本,保存到同名txt文件中
param:
filePath: 需要读取的pdf文档的目录
introduced module:
from pdfminer.pdfpage import PDFPage
from pdfminer.pdfparser import PDFParser
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.converter import PDFPageAggregator
from pdfminer.layout import LAParams
from pdfminer.pdfdocument import PDFDocument, PDFTextExtractionNotAllowed
import os.path
"""
file = open(filePath, 'rb')
praser = PDFParser(file)
doc = PDFDocument(praser, '')
praser.set_document(doc)
if not doc.is_extractable:
raise PDFTextExtractionNotAllowed
rsrcmgr = PDFResourceManager()
laparams = LAParams()
device = PDFPageAggregator(rsrcmgr, laparams=laparams)
interpreter = PDFPageInterpreter(rsrcmgr, device)
result = []
for page in PDFPage.create_pages(doc):
interpreter.process_page(page)
layout = device.get_result()
for x in layout:
if hasattr(x, "get_text"):
result.append(x.get_text())
fileNames = os.path.splitext(filePath)
with open(fileNames[0] + '.txt', 'a', encoding="utf-8") as f:
results = x.get_text()
f.write(results)