windows下Python实现将pdf文件转化为png格式图片的方法

2020-01-04 17:02:09

字体：大中小

来源：转载

供稿：网友

本文实例讲述了windows下Python实现将pdf文件转化为png格式图片的方法。分享给大家供大家参考，具体如下：

最近工作中需要把pdf文件转化为图片，想用Python来实现，于是在网上找啊找啊找啊找，找了半天，倒是找到一些代码。

1、第一个找到的代码，我试了一下好像是反了，只能实现把图片转为pdf，而不能把pdf转为图片。。。

参考链接：https://zhidao.baidu.com/question/745221795058982452.html

代码如下：

#!/usr/bin/env pythonimport osimport sysfrom reportlab.lib.pagesizes import A4, landscapefrom reportlab.pdfgen import canvasf = sys.argv[1]filename = ''.join(f.split('/')[-1:])[:-4]f_jpg = filename+'.jpg'print f_jpgdef conpdf(f_jpg): f_pdf = filename+'.pdf' (w, h) = landscape(A4) c = canvas.Canvas(f_pdf, pagesize = landscape(A4)) c.drawImage(f, 0, 0, w, h) c.save() print "okkkkkkkk."conpdf(f_jpg)

2、第二个是文章写的比较详细，可惜的是linux下的代码，所以仍然没用。

3、第三个文章指出有一个库PythonMagick可以实现这个功能，需要下载一个库 PythonMagick-0.9.10-cp27-none-win_amd64.whl 这个是64位的。

这里不得不说自己又犯了一个错误，因为自己从python官网上下载了一个python 2.7,以为是64位的版本，实际上是32位的版本，所以导致python的版本（32位）和下载的PythonMagick的版本（64位）不一致，弄到晚上12点多，总算了发现了这个问题。。。

4、然后，接下来继续用搜索引擎搜，找到很多stackoverflow的问题帖子，发现了2个代码，不过要先下载PyPDF2以及ghostscript模块。

先通过pip来安装 PyPDF2、PythonMagick、ghostscript 模块。

C:/Users/Administrator>pip install PyPDF2Collecting PyPDF2 Using cached PyPDF2-1.25.1.tar.gzInstalling collected packages: PyPDF2 Running setup.py install for PyPDF2Successfully installed PyPDF2-1.25.1You are using pip version 7.1.2, however version 8.1.2 is available.You should consider upgrading via the 'python -m pip install --upgrade pip' command.C:/Users/Administrator>pip install C:/PythonMagick-0.9.10-cp27-none-win_amd64.whlProcessing c:/pythonmagick-0.9.10-cp27-none-win_amd64.whlInstalling collected packages: PythonMagickSuccessfully installed PythonMagick-0.9.10You are using pip version 7.1.2, however version 8.1.2 is available.You should consider upgrading via the 'python -m pip install --upgrade pip' command.C:/Users/Administrator>pip install ghostscriptCollecting ghostscript Downloading ghostscript-0.4.1.tar.bz2Requirement already satisfied (use --upgrade to upgrade): setuptools in c:/python27/lib/site-packages (from ghostscript)Installing collected packages: ghostscript Running setup.py install for ghostscriptSuccessfully installed ghostscript-0.4.1You are using pip version 7.1.2, however version 8.1.2 is available.You should consider upgrading via the 'python -m pip install --upgrade pip' command.

下面是代码

代码1：

import osimport ghostscriptfrom PyPDF2 import PdfFileReader, PdfFileWriterfrom tempfile import NamedTemporaryFilefrom PythonMagick import Imagereader = PdfFileReader(open("C:/deep.pdf", "rb"))for page_num in xrange(reader.getNumPages()): writer = PdfFileWriter() writer.addPage(reader.getPage(page_num)) temp = NamedTemporaryFile(prefix=str(page_num), suffix=".pdf", delete=False) writer.write(temp) print temp.name tempname = temp.name temp.close() im = Image(tempname) #im.density("3000") # DPI, for better quality #im.read(tempname) im.write("some_%d.png" % (page_num)) os.remove(tempname)

代码2：

import sysimport PyPDF2import PythonMagickimport ghostscriptpdffilename = "C:/deep.pdf"pdf_im = PyPDF2.PdfFileReader(file(pdffilename, "rb"))print '1'npage = pdf_im.getNumPages()print('Converting %d pages.' % npage)for p in range(npage): im = PythonMagick.Image() im.density('300') im.read(pdffilename + '[' + str(p) +']') im.write('file_out-' + str(p)+ '.png') #print pdffilename + '[' + str(p) +']','file_out-' + str(p)+ '.png'

然后执行时都报错了，这个是代码2 的报错信息：

Traceback (most recent call last): File "C:/c.py", line 15, in <module> im.read(pdffilename + '[' + str(p) +']')RuntimeError: pythonw.exe: PostscriptDelegateFailed `C:/DEEP.pdf': No such file or directory @ error/pdf.c/ReadPDFImage/713

总是在上面的 im.read(pdffilename + '[' + str(p) +']') 这一行报错。

于是，根据报错的信息在网上查，但是没查到什么有用的信息，但是感觉应该和GhostScript有关，于是在网上去查安装包，找到一个在github上的下载连接，但是点进去的时候显示无法下载。

最后，在csdn的下载中找到了这个文件：GhostScript_Windows_9.15_win32_win64，安装了64位版本，之后，再次运行上面的代码，都能用了。

不过代码2需要做如下修改，不然还是会报 No such file or directory @ error/pdf.c/ReadPDFImage/713 错误：

#代码2import sysimport PyPDF2import PythonMagickimport ghostscriptpdffilename = "C:/deep.pdf"pdf_im = PyPDF2.PdfFileReader(file(pdffilename, "rb"))print '1'npage = pdf_im.getNumPages()print('Converting %d pages.' % npage)for p in range(npage): im = PythonMagick.Image(pdffilename + '[' + str(p) +']') im.density('300') #im.read(pdffilename + '[' + str(p) +']') im.write('file_out-' + str(p)+ '.png') #print pdffilename + '[' + str(p) +']','file_out-' + str(p)+ '.png'

这次有个很深刻的体会，就是解决这个问题过程中，大部分时间都是用在查资料、验证资格资料是否有用上了，搜索资料的能力很重要。

而在实际搜索资料的过程中，国内关于PythonMagick的文章太少了，搜索出来的大部分有帮助的文章都是国外的，但是这些国外的帖子文章，也没有解决我的问题或者是给出有用的线索，最后还是通过自己的思考，解决了问题。

希望本文所述对大家Python程序设计有所帮助。

上一篇：python僵尸进程产生的原因

下一篇：Python随机生成手机号、数字的方法详解