在 PDF 中,您可以更改页面大小以使文档满足不同的需求。例如,在创建讲义或压缩版本的文档时需要较小的页面大小,而较大的页面大小可能对设计海报或图形密集型材料有用。在某些情况下,您可能还需要获取页面尺寸(宽度和高度)以确定文档是否进行了最佳调整。本文将介绍如何使用 Spire.PDF for Python 在 Python 中以编程方式更改或获取 PDF 页面大小。
本教程需要用到 Spire.PDF for Python 和 plum-dispatch v1.7.4。可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.PDF如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.PDF for Python
更改 PDF 文件页面大小的方法是创建一个新的 PDF 文件,并添加所需大小的页面,然后根据原始 PDF 文件中的页面创建模板,并将模板绘制到新 PDF 文件中的页面上。此过程将保留原始 PDF 文件中的文本、图像和其他元素。
Spire.PDF for Python 支持各种标准纸张尺寸,如 letter、legal、A0、A1、A2、A3、A4、B0、B1、B2、B3、B4 等。以下是将 PDF 文件的页面尺寸更改为标准纸张尺寸的步骤:
from spire.pdf.common import *
from spire.pdf import *
# 设置输入文件路径,示例文档.pdf为待处理的PDF文件
inputFile = "示例文档.pdf"
# 设置输出文件路径,修改页面大小后保存的PDF文件
outputFile = "修改页面大小为B4.pdf"
# 创建原始PdfDocument对象
originalPdf = PdfDocument()
# 加载待处理的PDF文件到原始PDF文档对象中
originalPdf.LoadFromFile(inputFile)
# 创建新的PdfDocument对象用于存储修改后的结果
newPdf = PdfDocument()
# 遍历原始PDF文档中的每一页
for i in range(originalPdf.Pages.Count):
# 获取当前页对象
page = originalPdf.Pages.get_Item(i)
# 创建新的页面,并设置页面大小为B4,边距为0
newPage = newPdf.Pages.Add(PdfPageSize.B4(), PdfMargins(0.0))
# 创建文本布局对象
layout = PdfTextLayout()
# 设置文本布局类型为单页
layout.Layout = PdfLayoutType.OnePage
# 创建模板对象
template = page.CreateTemplate()
# 在新页面上绘制模板内容,并应用文本布局
template.Draw(newPage, PointF.Empty(), layout)
# 将修改后的PDF文档保存到指定路径
newPdf.SaveToFile(outputFile)
newPdf.Close()
originalPdf.Close()
Spire.PDF for Python 使用磅(1/72英寸)作为度量单位。如果你需要将 PDF 的页面大小更改为其他度量单位(如英寸或毫米)的自定义纸张大小,可以使用 PdfUnitConvertor 类将其转换为点。
以下是将 PDF 文件的页面大小更改为自定义纸张大小的步骤,以英寸为单位:
from spire.pdf.common import *
from spire.pdf import *
# 设置输入文件路径,示例文档.pdf为待处理的PDF文件
inputFile = "示例文档.pdf"
# 设置输出文件路径,自定义页面大小后保存的PDF文件
outputFile = "自定义页面大小.pdf"
# 创建原始PdfDocument对象
originalPdf = PdfDocument()
# 加载待处理的PDF文件到原始PDF文档对象中
originalPdf.LoadFromFile(inputFile)
# 创建新的PdfDocument对象用于存储修改后的结果
newPdf = PdfDocument()
# 创建单位转换器对象
unitCvtr = PdfUnitConvertor()
# 将宽度从12.0英寸转换为磅(Point)单位
width = unitCvtr.ConvertUnits(12.0, PdfGraphicsUnit.Inch, PdfGraphicsUnit.Point)
# 将高度从15.5英寸转换为磅(Point)单位
height = unitCvtr.ConvertUnits(15.5, PdfGraphicsUnit.Inch, PdfGraphicsUnit.Point)
# 创建自定义大小
size = SizeF(width, height)
# 遍历原始PDF文档中的每一页
for i in range(originalPdf.Pages.Count):
# 获取当前页对象
page = originalPdf.Pages.get_Item(i)
# 创建新的页面,并设置页面大小为自定义大小,边距为0
newPage = newPdf.Pages.Add(size, PdfMargins(0.0))
# 创建文本布局对象
layout = PdfTextLayout()
# 设置文本布局类型为单页
layout.Layout = PdfLayoutType.OnePage
# 创建模板对象
template = page.CreateTemplate()
# 在新页面上绘制模板内容,并应用文本布局
template.Draw(newPage, PointF.Empty(), layout)
# 将修改后的PDF文档保存到指定路径
newPdf.SaveToFile(outputFile)
newPdf.Close()
originalPdf.Close()
Spire.PDF for Python 提供了 PdfPageBase.Size.Width 和 PdfPageBase.Size.Height 属性来获取 PDF 页面的宽度和高度(以磅为单位)。如果您想将默认的度量单位转换为其他单位,可以使用 PdfUnitConvertor 类。
下面是获取 PDF 页面大小的步骤:
from spire.pdf.common import *
from spire.pdf import *
# 定义函数:将文本内容追加到文件中
def AppendAllText(fname: str, text: List[str]):
fp = open(fname, "w")
for s in text:
fp.write(s + "\n")
fp.close()
# 设置输入文件路径,示例文档.pdf为待处理的PDF文件
inputFile = "示例文档.pdf"
# 设置输出文件路径,获取页面大小后保存的文本文件
outputFile = "获取页面大小.txt"
# 创建PdfDocument对象
pdf = PdfDocument()
# 加载待处理的PDF文件到PDF文档对象中
pdf.LoadFromFile(inputFile)
# 获取第一页对象
page = pdf.Pages.get_Item(0)
# 获取页面宽度和高度(以磅为单位)
pointWidth = page.Size.Width
pointHeight = page.Size.Height
# 创建单位转换器对象
unitCvtr = PdfUnitConvertor()
# 将页面宽度和高度从磅(Point)转换为像素(Pixel)单位
pixelWidth = unitCvtr.ConvertUnits(pointWidth, PdfGraphicsUnit.Point, PdfGraphicsUnit.Pixel)
pixelHeight = unitCvtr.ConvertUnits(pointHeight, PdfGraphicsUnit.Point, PdfGraphicsUnit.Pixel)
# 将页面宽度和高度从磅(Point)转换为英寸(Inch)单位
inchWidth = unitCvtr.ConvertUnits(pointWidth, PdfGraphicsUnit.Point, PdfGraphicsUnit.Inch)
inchHeight = unitCvtr.ConvertUnits(pointHeight, PdfGraphicsUnit.Point, PdfGraphicsUnit.Inch)
# 将页面宽度和高度从磅(Point)转换为厘米(Centimeter)单位
centimeterWidth = unitCvtr.ConvertUnits(pointWidth, PdfGraphicsUnit.Point, PdfGraphicsUnit.Centimeter)
centimeterHeight = unitCvtr.ConvertUnits(pointHeight, PdfGraphicsUnit.Point, PdfGraphicsUnit.Centimeter)
# 创建用于保存内容的列表
content = []
# 将页面大小信息添加到内容列表中
content.append("文件的页面大小(以磅为单位)为(宽度: " + str(pointWidth) + "pt, 高度: " + str(pointHeight) + "pt).")
content.append("文件的页面大小(以像素为单位)为(宽度: " + str(pixelWidth) + "pixel, 高度: " + str(pixelHeight) + "pixel).")
content.append("文件的页面大小(以英寸为单位)为(宽度: " + str(inchWidth) + "inch, 高度: " + str(inchHeight) + "inch).")
content.append("文件的页面大小(以厘米为单位)为(宽度: " + str(centimeterWidth) + "cm, 高度: " + str(centimeterHeight) + "cm.)")
# 调用函数将内容写入输出文件中
AppendAllText(outputFile, content)
pdf.Close()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
在 Excel 中,对行和列进行分组可以提供更有组织和结构化的数据视图,从而更容易分析和理解复杂的数据集。通过分组行或列,您可以创建一个层次结构,将相关的数据分组在一起,并能够展开或折叠特定部分。这样可以以汇总视图显示数据,同时隐藏详细信息。在本文中,您将学习如何使用 Spire.XLS for Python 对行和列进行分组或取消分组,以及如何折叠或展开 Excel 分组。
此教程需要 Spire.XLS for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.XLS如果您不确定如何安装,请参考: 如何在 Windows 中安装 Spire.XLS for Python
Spire.XLS for Python 提供的 Worksheet.GroupByRows() 和 Worksheet.GroupByColumns() 方法可用于对 Excel 工作表中的特定行和列进行分组。具体步骤如下:
from spire.xls import *
from spire.xls.common import *
inputFile = "示例.xlsx"
outputFile = "Excel分组.xlsx"
# 创建Workbook对象
workbook = Workbook()
# 加载Excel文件
workbook.LoadFromFile(inputFile)
# 获取第一张工作表
sheet = workbook.Worksheets[0]
# 对指定行进行分组
sheet.GroupByRows(2, 6, False)
sheet.GroupByRows(8, 14, False)
# 对指定列进行分组
sheet.GroupByColumns(5, 8, False)
# 保存结果文件
workbook.SaveToFile(outputFile, ExcelVersion.Version2016)
workbook.Dispose()
要取消 Excel 工作表中行和列的分组,可以使用 Worksheet.UngroupByRows() 和 Worksheet.UngroupByColumns() 方法。具体步骤如下:
from spire.xls import *
from spire.xls.common import *
inputFile = "Excel分组.xlsx"
outputFile = "取消分组.xlsx"
# 创建Workbook对象
workbook = Workbook()
# 加载Excel文件
workbook.LoadFromFile(inputFile)
# 获取第一张工作表
sheet = workbook.Worksheets[0]
# 取消行和列的分组
sheet.UngroupByRows(2, 6)
sheet.UngroupByRows(8, 14)
sheet.UngroupByColumns(5, 8)
# 保存结果文件
workbook.SaveToFile(outputFile, ExcelVersion.Version2016)
workbook.Dispose()
Excel 中的展开或折叠分组是指显示或隐藏分组部分中的详细信息。使用 Spire.XLS for Python,可以通过 Worksheet.Range[].ExpandGroup() 或 Worksheet.Range[].CollapseGroup() 方法展开或折叠分组。具体步骤如下:
from spire.xls import *
from spire.xls.common import *
inputFile = "分组.xlsx"
outputFile = "展开或折叠分组.xlsx"
# 创建Workbook对象
workbook = Workbook()
# 加载Excel文件
workbook.LoadFromFile(inputFile)
# 获取第一张工作表
sheet = workbook.Worksheets[0]
# 展开分组
sheet.Range["A2:H6"].ExpandGroup(GroupByType.ByRows)
# 折叠分组
sheet.Range["E1:H16"].CollapseGroup(GroupByType.ByColumns)
# 保存结果文件
workbook.SaveToFile(outputFile, ExcelVersion.Version2016)
workbook.Dispose()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
SVG(可缩放矢量图形)是一种基于 XML 的图形格式,它可以无损地缩放而不失去清晰度和质量,因此在 Web 图形和矢量插图领域得到了广泛地应用。而 PDF 则是一种通用的文件格式,在各种设备和操作系统上广受支持。将 SVG 文件转换为 PDF 格式能够方便文件共享,确保接收者可以轻松打开和查看文件,而无需安装特定软件或担心浏览器兼容性问题。 这篇文章将介绍如何使用 Python 和 Spire.PDF for Python 库将 SVG 文件转换为 PDF 格式。
本教程需要用到 Spire.PDF for Python 和 plum-dispatch v1.7.4。可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.PDF如果您不清楚如何安装,请参考此教程: 如何在 Windows 中安装 Spire.PDF for Python
Spire.PDF for Python 为用户提供了 PdfDocument.LoadFromSvg() 方法,用于加载 SVG 文件。加载后,用户可以使用 PdfDocument.SaveToFile() 方法轻松将 SVG 文件保存为 PDF 格式。详细步骤如下:
from spire.pdf.common import *
from spire.pdf import *
# 创建PdfDocument类的对象
doc = PdfDocument()
# 加载SVG文件
doc.LoadFromSvg("https://cdn.e-iceblue.cn/Sample.svg")
# 将该SVG文件保存为PDF格式
doc.SaveToFile("ConvertSvgToPdf.pdf", FileFormat.PDF)
doc.Close()
除了直接将 SVG 转换为 PDF 之外,Spire.PDF for Python 还支持将 SVG 文件添加到 PDF 的特定位置。详细步骤如下:
from spire.pdf.common import *
from spire.pdf import *
# 创建PdfDocument类的对象
doc1 = PdfDocument()
# 加载SVG文件
doc1.LoadFromSvg("https://cdn.e-iceblue.cn/Sample.svg")
# 根据SVG文件的内容创建一个模板
template = doc1.Pages[0].CreateTemplate()
# 获取模板的宽度和高度
width = template.Width
height = template.Height
# 创建另一个PdfDocument类的对象
doc2 = PdfDocument()
# 加载PDF文档
doc2.LoadFromFile("Sample.pdf")
# 将模板绘制到PDF第一页的指定位置
doc2.Pages[0].Canvas.DrawTemplate(template, PointF(10.0, 100.0), SizeF(width*0.8, height*0.8))
# 保存结果文档
doc2.SaveToFile("AddSvgToPdf.pdf", FileFormat.PDF)
doc2.Close()
doc1.Close()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
通过将 HTML 转换为图像,可以捕捉 HTML 内容的外观和布局,并生成静态图像文件。这一技术具有广泛的应用领域,包括生成网站预览、创建屏幕截图、对网页进行归档、以及将 HTML 内容整合到主要处理图像的应用程序中。本文将介绍在 Python 环境中利用 Spire.Doc for Python 库将 HTML 文件或 HTML 字符串转换为图像的方法。
本教程需要用到 Spire.Doc for Python 和 plum-dispatch v1.7.4。可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.Doc如果您不确定如何安装,请参考:如何在 Windows 中安装 Spire.Doc for Python
在使用 Document.LoadFromFile() 方法将 HTML 文件加载为 Document 对象时,该方法会自动将其内容呈现为 Word 文档的形式。然后,开发者可以使用 Document.SaveImageToStreams() 方法将指定页面保存为图像流并写入图像文件。
以下是使用 Python 将 HTML 文件转换为图像的操作步骤:
from spire.doc import Document
from spire.doc import FileFormat
from spire.doc import XHTMLValidationType
from spire.doc import ImageType
# 创建Document类的对象
doc = Document()
# 载入HTML文件
doc.LoadFromFile("G:/文档/示例22.html", FileFormat.Html, XHTMLValidationType.none)
# 将文档第二页保存为图像
imageStream = doc.SaveImageToStreams(0, ImageType.Bitmap)
# 将图像写入PNG文件
with open("output/HTML转图像.png", "wb") as imageFile:
imageFile.write(imageStream.ToArray())
doc.Close()
使用 Paragraph.AppendHTML() 方法可以将简单的 HTML 字符串(通常是文本、链接及其格式)呈现为 Word 文档内容,然后再使用 Document.SaveImageToStreams() 方法将其转换为图像流并写入文件,即可完成 HTML 字符串到图像的转换。
以下是在 Python 程序中将 HTML 字符串转换为图像的操作步骤:
from spire.doc import Document
from spire.doc import ImageType
# 创建Document类的对象
doc = Document()
# 添加一个节到文档
sec = doc.AddSection()
# 添加一个段落到节
par = sec.AddParagraph()
# 指定HTML字符串
htmlString = """
<html>
<head>
<title>旅游统计信息</title>
<style>
h1 {
color: blue;
font-family: Arial, sans-serif;
}
p {
color: green;
font-family: Verdana, sans-serif;
}
ul {
color: red;
font-family: "Courier New", monospace;
}
ol {
color: purple;
font-family: "Times New Roman", serif;
}
table {
border-collapse: collapse;
width: 100%;
}
th, td {
border: 1px solid black;
padding: 8px;
text-align: left;
font-family: Arial, sans-serif;
}
th {
background-color: #f2f2f2;
}
</style>
</head>
<body>
<h1>旅游统计信息</h1>
<p>以下是一些旅游方面的统计信息:</p>
<h2>热门旅游目的地</h2>
<ul>
<li>巴黎</li>
<li>罗马</li>
<li>东京</li>
<li>纽约</li>
</ul>
<h2>旅游收入排行榜</h2>
<table>
<tr>
<th>国家</th>
<th>收入(亿美元)</th>
</tr>
<tr>
<td>法国</td>
<td>79.5</td>
</tr>
<tr>
<td>美国</td>
<td>76.9</td>
</tr>
<tr>
<td>西班牙</td>
<td>67.7</td>
</tr>
<tr>
<td>意大利</td>
<td>58.3</td>
</tr>
</table>
<a href="https://www.example.com">点击这里访问示例网站</a>
</body>
</html>
"""
# 将HTML字符串所展示的内容插入到段落中
par.AppendHTML(htmlString)
# 将文档第一页转换为图像
imageStream = doc.SaveImageToStreams(0, ImageType.Bitmap)
# 将图像写入PNG文件
with open("output/HTML字符串转图像.png", "wb") as imageFile:
imageFile.write(imageStream.ToArray())
doc.Close()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。

将 HTML 转换为 PDF 是一项常见的需求,尤其是使用 Python 生成可打印的报表、保存网页内容,或创建格式统一的离线文档时。为了避免手动操作的繁琐,本文将带你了解如何使用 Python 将 HTML 转换为 PDF,无论是本地 HTML 文件,还是 HTML 字符串。如果你正在寻找一种简单可靠的方式在 Python 中生成 PDF 文件,这篇教程将为你提供实用且有效的指导。
想要在 Python 中轻松将 HTML 转换为 PDF,首先需要一个可靠的库,用于处理 HTML 解析并生成 PDF。本文采用的 Spire.Doc for Python 是一款功能强大且易于使用的 HTML 到 PDF 转换库,可直接将 HTML 内容转换为 PDF,无需依赖浏览器、无头引擎或第三方工具。
pip install spire.doc
同时,Spire.Doc 还提供免费版本,适用于小型项目或评估。
在安装完成必要的 Python 库后,你只需几行代码即可在 Python 中实现将 HTML 保存为 PDF。
我们先来看看最常见的任务,将 HTML 文件直接转换为 PDF。在 Spire.Doc 的帮助下,首先通过Document.LoadFromFile() 方法加载 .html 文件。然后调用 Document.SaveToFile() 方法即可将其保存为 PDF。
使用 Python 将 HTML 文件转换为 PDF 的步骤:
下方代码演示了如何在 Python 中将单个 HTML 文件直接转换为 PDF:
from spire.doc import Document
from spire.doc import FileFormat
from spire.doc import XHTMLValidationType
# 创建 Document 类的对象
doc = Document()
# 载入 HTML 文件
doc.LoadFromFile("/input/百科全书.html", FileFormat.Html, XHTMLValidationType.none)
# 将 HTML 文件转换为 PDF 文件并保存
doc.SaveToFile("/output/HTML转PDF.pdf", FileFormat.PDF)
doc.Close()
转换结果文档预览:

如果你希望将 HTML 字符串转换为 PDF,Spire.Doc 同样提供简单直接的解决方案。对于包含段落、文本样式及基础格式的 HTML 内容,可以使用 Paragraph.AppendHTML() 方法将 HTML 插入到 Word 文档中,随后通过 Document.SaveToFile() 方法将文档保存为 PDF。
将 HTML 字符串转换为 PDF 的详细步骤如下:
下方是完整的 Python 示例代码,演示如何将 HTML 字符串转换为 PDF:
from spire.doc import Document
from spire.doc import FileFormat
# 创建 Document 类的对象
doc = Document()
# 添加一个节到文档
sec = doc.AddSection()
# 添加一个段落到节
par = sec.AddParagraph()
# 指定 HTML 字符串
htmlString = """
<html>
<head>
<title>旅游统计信息</title>
<style>
h1 {
color: blue;
font-family: Arial, sans-serif;
}
p {
color: green;
font-family: Verdana, sans-serif;
}
ul {
color: red;
font-family: "Courier New", monospace;
}
ol {
color: purple;
font-family: "Times New Roman", serif;
}
table {
border-collapse: collapse;
width: 100%;
}
th, td {
border: 1px solid black;
padding: 8px;
text-align: left;
font-family: Arial, sans-serif;
}
th {
background-color: #f2f2f2;
}
</style>
</head>
<body>
<h1>旅游统计信息</h1>
<p>以下是一些旅游方面的统计信息:</p>
<h2>热门旅游目的地</h2>
<ul>
<li>巴黎</li>
<li>罗马</li>
<li>东京</li>
<li>纽约</li>
</ul>
<h2>旅游收入排行榜</h2>
<table>
<tr>
<th>国家</th>
<th>收入(亿美元)</th>
</tr>
<tr>
<td>法国</td>
<td>79.5</td>
</tr>
<tr>
<td>美国</td>
<td>76.9</td>
</tr>
<tr>
<td>西班牙</td>
<td>67.7</td>
</tr>
<tr>
<td>意大利</td>
<td>58.3</td>
</tr>
</table>
<a href="https://www.example.com">点击这里访问示例网站</a>
</body>
</html>
"""
# 将 HTML 字符串所展示的内容插入到段落中
par.AppendHTML(htmlString)
# 将文档转换为 PDF 文件并保存
doc.SaveToFile("/output/HTML字符串转PDF.pdf", FileFormat.PDF)
doc.Close()
HTML 字符串转换为 PDF 的结果预览:

虽然在 Python 中将 HTML 转换为 PDF 显得简单直接,但有时你可能希望对输出结果进行更精细的控制。例如,你可能需要为 PDF 添加密码以保护文档,或者嵌入字体以确保在不同设备上呈现一致的格式效果。
在本节中,你将学习如何使用 Spire.Doc 自定义 HTML 转 PDF 的转换过程。
为了防止未经授权的查看或编辑,你可以通过设置用户密码和所有者密码对生成的 PDF 文档进行加密。
from spire.doc import Document, FileFormat, XHTMLValidationType, ToPdfParameterList, PdfPermissionsFlags, PdfEncryptionKeySize
# 创建一个 ToPdfParameterList 对象
toPdf = ToPdfParameterList()
# 设置密码
userPassword = "viewer"
ownerPassword = "E-iceblue"
toPdf.PdfSecurity.Encrypt(userPassword, ownerPassword, PdfPermissionsFlags.Default, PdfEncryptionKeySize.Key128Bit)
为了确保生成的 PDF 在所有设备上都能正确显示,你可以将文档中使用的所有字体嵌入到 PDF 文件中,从而避免因缺少字体而导致的排版错乱问题。
from spire.doc import Document, FileFormat, XHTMLValidationType, ToPdfParameterList
# 创建一个 ToPdfParameterList 对象,并设置嵌入字体
ppl = ToPdfParameterList()
ppl.IsEmbeddedAllFonts = True
这些选项可以更加细致地控制 HTML 转 PDF 的输出文件,尤其适用于专业文档共享或长期存储等场景。
借助 Spire.Doc for Python,将 HTML 转换为 PDF 变得简单且灵活。无论是处理静态 HTML 文件,还是动态生成的 HTML 字符串,亦或是需要对 PDF 进行加密和个性化设置,该库都能提供完整的解决方案 —— 所需代码仅需几行。
感兴趣的话,不妨该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。,立即获取免费 30 天授权,开始在 Python 中高效生成高质量的 PDF 文档!
Q1:我可以在 Python 中将 HTML 文件转换为 PDF 吗?
可以。使用 Spire.Doc for Python,只需几行代码即可将本地 HTML 文件转换为 PDF。
Q2:如何在 Chrome 中将 HTML 转换为 PDF?
虽然 Chrome 支持手动“另存为 PDF”,但不适合批量或自动化处理。如果你使用 Python,Spire.Doc 提供了更高效的编程方式来完成 HTML 到 PDF 的转换。
Q3:如何在转换过程中保持 HTML 格式不丢失?
为了最大程度保留格式,请注意以下几点:
Spire.PDF for Java 10.1.3 已发布。本次更新新增 PdfTextReplacer 接口来实现替换文本的功能以及 PdfImageHelper 接口来实现删除图片、提取图片、替换图片和压缩图片的功能。此外,本次更新还提高了绘制水印的效率。详情请阅读以下内容。
新功能:
PdfDocument pdf = new PdfDocument();
pdf.loadFromFile("sample.pdf");
PdfPageBase page = pdf.getPages().get(0);
PdfTextReplacer replacer = new PdfTextReplacer(page);
PdfTextReplaceOptions options= new PdfTextReplaceOptions();
options.setReplaceType(EnumSet.of(ReplaceActionType.WholeWord));
replacer.replaceText("www.google.com", "1234567");
pdf.saveToFile(outputFile);PdfImageHelper imageHelper = new PdfImageHelper();
PdfImageInfo[] imageInfoCollection= imageHelper.getImagesInfo(page);
Delete image:
imageHelper.deleteImage(imageInfoCollection[0]);
Extract image:
int index = 0;
for (com.spire.pdf.utilities.PdfImageInfo img : imageInfoCollection) {
BufferedImage image = img.getImage();
File output = new File(outputFile_Img + String.format("img_%d.png", index));
ImageIO.write(image, "PNG", output);
index++;
}PdfImage image = PdfImage.fromFile("ImgFiles/E-iceblue logo.png");
imageHelper.replaceImage(imageInfoCollection[i], image);
Compress image:
for (PdfPageBase page : (Iterable<PdfPageBase>)doc.getPages())
{
if (page != null)
{
if (imageHelper.getImagesInfo(page) != null)
{
for (com.spire.pdf.utilities.PdfImageInfo info : imageHelper.getImagesInfo(page))
{
info.tryCompressImage();
}
}
}
}问题修复:
Spire.PDF 10.1已发布。该版本增强了.NET Standard平台上从PDF到图片的转换功能。此外,还修复了一系列其他已知问题,例如打印PDF时内容显示不清晰的问题。详情请阅读以下内容。
问题修复:
Spire.XLS 14.1 已发布。本次更新在 FileFormat 枚举中增加了 XLT 、XLTX、 XLTM 文档格式,同时改善了转换工作表到图片时占用的内存量。此外,本次更新还增强了 Excel 到 PDF 和 CSV 的转换。一些已知问题也在该版本中得到修复,如了获取单元格失败的问题。详情请阅读以下内容。
新功能:
问题修复:
PDF 书签是一种导航辅助工具,它允许用户快速定位并跳转到 PDF 文档中的特定章节或页面。通过简单地点击书签,用户可以直达目标位置,无需手动滚动或搜索冗长文档中的内容。本文将介绍如何使用 Spire.PDF for Python 以编程方式添加、修改和删除 PDF 中的书签。
本教程需要 Spire.PDF for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.PDF如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.PDF for Python
Spire.PDF for Python 提供了向 PDF 文档添加书签的方法: PdfDocument.Bookmarks.Add()。您可以使用此方法为 PDF 文档创建主要书签,并使用 PdfBookmarkCollection.Add() 方法为主要书签添加子书签。此外,PdfBookmark 类还提供了其他方法来设置书签的属性,如目标位置、文本颜色和文本样式。以下是向 PDF 文档添加书签的详细步骤。
from spire.pdf.common import *
from spire.pdf import *
from spire.pdf.common import *
from spire.pdf import *
# 创建 PdfDocument 对象
doc = PdfDocument()
# 加载 PDF 文件
doc.LoadFromFile("示例.pdf")
# 遍历 PDF 文件中的页面
for i in range(doc.Pages.Count):
page = doc.Pages.get_Item(i)
# 设置书签的标题和目标位置
bookmarkTitle = "书签-{0}".format(i+1)
bookmarkDest = PdfDestination(page, PointF(0.0, 0.0))
# 创建并配置书签
bookmark = doc.Bookmarks.Add(bookmarkTitle)
bookmark.Color = PdfRGBColor(Color.get_SaddleBrown())
bookmark.DisplayStyle = PdfTextStyle.Bold
bookmark.Action = PdfGoToAction(bookmarkDest)
# 创建集合以容纳子书签
bookmarkColletion = PdfBookmarkCollection(bookmark)
# 设置子书签的标题和目标位置
childBookmarkTitle = "子书签-{0}".format(i+1)
childBookmarkDest = PdfDestination(page, PointF(0.0, 100.0))
# 创建并配置子书签
childBookmark = bookmarkColletion.Add(childBookmarkTitle)
childBookmark.Color = PdfRGBColor(Color.get_Coral())
childBookmark.DisplayStyle = PdfTextStyle.Italic
childBookmark.Action = PdfGoToAction(childBookmarkDest)
# 保存 PDF 文件
outputFile = "书签.pdf"
doc.SaveToFile(outputFile)
# 关闭文档
doc.Close()
如果您需要更新现有的书签,可以使用 PdfBookmark 类的方法来重命名书签并更改其文本颜色、文本样式。以下是详细的步骤:
from spire.pdf.common import *
from spire.pdf import *
# 创建一个 PdfDocument 对象
doc = PdfDocument()
# 加载一个 PDF 文件
doc.LoadFromFile("书签.pdf")
# 获取第一个书签
bookmark = doc.Bookmarks.get_Item(0)
# 修改书签的标题
bookmark.Title = "被修改的书签"
# 设置书签的颜色
bookmark.Color = PdfRGBColor(Color.get_Black())
# 设置书签的文本样式
bookmark.DisplayStyle = PdfTextStyle.Bold
# 编辑父书签下的子书签
pBookmark = PdfBookmarkCollection(bookmark)
for i in range(bookmark.Count):
childBookmark = pBookmark.get_Item(i)
childBookmark.Color = PdfRGBColor(Color.get_Blue())
childBookmark.DisplayStyle = PdfTextStyle.Regular
# 保存 PDF 文档
outputFile = "修改书签.pdf"
# 关闭文档
doc.SaveToFile(outputFile)
Spire.PDF for Python 还提供了删除 PDF 文档中任何书签的方法。PdfDocument.Bookmarks.RemoveAt() 方法用于删除特定的主要书签,PdfDocument.Bookmarks.Clear() 方法用于删除所有书签,而 PdfBookmarkCollection.RemoveAt() 方法用于删除主要书签的特定子书签。从 PDF 文档中删除书签的详细步骤如下:
from spire.pdf.common import *
from spire.pdf import *
# 创建一个 PdfDocument 对象
doc = PdfDocument()
# 加载一个 PDF 文件
doc.LoadFromFile("书签.pdf")
# # 删除第一个书签
# doc.Bookmarks.RemoveAt(0)
# # 获取第一个书签
# bookmark = doc.Bookmarks[0]
# # 从第一个书签中删除第一个子书签
# pBookmark = PdfBookmarkCollection(bookmark)
# pBookmark.RemoveAt(0)
# 删除所有书签
doc.Bookmarks.Clear()
# 保存 PDF 文档
output = "删除所有书签.pdf"
doc.SaveToFile(output)
# 关闭文档
doc.Close()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
Spire.Doc for Java 12.1.0已发布。本次更新移除了对Spire.Pdf.jar的依赖,且将应用授权的方法更改为com.spire.doc.license.LicenseProvider.setLicenseKey(key)。此外,还新增了一系列新功能,如新增添加图片水印的方法。详情请阅读以下内容。
调整:
新功能:
Document构造中设置newEngine不再起作用,内部默认采用新引擎
HeaderType枚举
GroupedShapeCollection类
ShapeObjectTextCollection类
MailMergeData接口
EnumInterface接口
public PictureWaterMark(InputStream inputeStream,boolean washout)
public PictureWaterMark(String filename,boolean washout)
Field类中downloadImage方法
IDocOleObject接口
PointsConverter类com.spire.license.LicenseProvider -> com.spire.doc.License.LicenseProvider// 设置自定义字体
Document.setCustomFontsFolders(string filePath);
// 处理自定义字体
Document.clearCustomFontsFolders();
// 清除缓存中占用内存的系统字体缓存
Document.clearSystemFontCache();
Example code:
Document doc = new Document();
doc.loadFromFile("inputFile.docx");
doc.setCustomFontsFolders(@"d:\Fonts");
doc.saveToFile("output.pdf", FileFormat.PDF);
doc.close();
doc.dispose();com.spire.doc.FileFormat.WPS -> com.spire.doc.FileFormat.Wps
com.spire.doc.FileFormat.WPT -> com.spire.doc.FileFormat.Wpt
ComparisonLevel -> TextDiffModeComparisonLevel getLevel() -> getTextCompareLevel()
setLevel(ComparisonLevel value) -> setTextCompareLevel(TextDiffMode)
IsPasswordProtect() -> isEncrypted()
getFillEfects() -> getFillEffects()File imageFile = new File("data/E-iceblue.png");
BufferedImage bufferedImage = ImageIO.read(imageFile);
// 通过输入 BufferedImage 创建 PictureWatermark 类的新实例,并设置水印图像的缩放因子
PictureWatermark picture = new PictureWatermark(bufferedImage,false);
// 或者使用另一种创建 PictureWatermark 的方法
// PictureWatermark picture = new PictureWatermark();
// picture.setPicture(bufferedImage);
// picture.isWashout(false);
// 设置水印图片的缩放比例
picture.setScaling(250);
// 设置要应用于文档的水印
document.setWatermark(picture);// 创建一个新的 Document 实例
Document document = new Document();
// 向文档添加一个节
Section section = document.addSection();
// 将段落添加到该节并向其附加文本
section.addParagraph().appendText("Line chart.");
// 添加一个新段落到该节
Paragraph newPara = section.addParagraph();
// 将折线图形状附加到指定宽度和高度的段落
ShapeObject shape = newPara.appendChart(ChartType.Line, 500, 300);
// 从形状中获取图表对象
Chart chart = shape.getChart();
// 获取图表的标题
ChartTitle title = chart.getTitle();
// 设置图表标题的文本
title.setText("My Chart");
// 清除图表中任何现有的系列
ChartSeriesCollection seriesColl = chart.getSeries();
seriesColl.clear();
// 定义类别(X 轴值)
String[] categories = { "C1", "C2", "C3", "C4", "C5", "C6" };
// 将两个具有指定类别和 Y 轴值的系列添加到图表中
seriesColl.add("AW Series 1", categories, new double[] { 1, 2, 2.5, 4, 5, 6 });
seriesColl.add("AW Series 2", categories, new double[] { 2, 3, 3.5, 6, 6.5, 7 });
// 将文档保存为Docx格式的文件
document.saveToFile("AppendLineChart.docx", FileFormat.Docx_2016);
// 释放文档资源
document.dispose();// 创建一个新的 Document 实例
Document doc = new Document();
// 从指定文件加载文档
doc.loadFromFile(inputFile);
// 使用加载的文档创建一个FixedLayoutDocument对象
FixedLayoutDocument layoutDoc = new FixedLayoutDocument(doc);
// 创建一个StringBuilder来存储提取的文本
StringBuilder stringBuilder = new StringBuilder();
// 获取第一页的第一行并将其附加到 StringBuilder
FixedLayoutLine line = layoutDoc.getPages().get(0).getColumns().get(0).getLines().get(0);
stringBuilder.append("Line: " + line.getText() + "\r\n");
// 检索与该行关联的原始段落并将其文本附加到 StringBuilder
Paragraph para = line.getParagraph();
stringBuilder.append("Paragraph text: " + para.getText() + "\r\n");
// 检索第一页上的所有文本,包括页眉和页脚,并将其附加到 StringBuilder
String pageText = layoutDoc.getPages().get(0).getText();
stringBuilder.append(pageText + "\r\n");
// 遍历文档中的每一页并打印每页的行数
for (Object obj : layoutDoc.getPages()) {
FixedLayoutPage page = (FixedLayoutPage) obj;
LayoutCollection<LayoutElement> lines = page.getChildEntities(LayoutElementType.Line, true);
stringBuilder.append("Page " + page.getPageIndex() + " has " + lines.getCount() + " lines." + "\r\n");
}
// 对第一段的布局实体执行反向查找并将它们附加到 StringBuilder
stringBuilder.append("\r\n");
stringBuilder.append("The lines of the first paragraph:" + "\r\n");
for (Object object : layoutDoc.getLayoutEntitiesOfNode(((Section) doc.getFirstChild()).getBody().getParagraphs().get(0))) {
FixedLayoutLine paragraphLine = (FixedLayoutLine) object;
stringBuilder.append(paragraphLine.getText().trim() + "\r\n");
stringBuilder.append(paragraphLine.getRectangle().toString() + "\r\n");
stringBuilder.append("");
}
// 将提取的文本写入文件
FileWriter fileWriter = new FileWriter(new File(outputFile));
fileWriter.write(stringBuilder.toString());
fileWriter.flush();
fileWriter.close();
// 释放文档资源
doc.close();
doc.dispose();// 创建一个新的文档对象
Document document = new Document();
// 在文档中添加一个新的节
Section section = document.addSection();
// 添加一个新的段落到该节
Paragraph paragraph = section.addParagraph();
// 将图片 (SVG) 附加到段落中
paragraph.appendPicture(inputSvg);
// 将文档保存到指定的输出文件
document.saveToFile(outputFile, FileFormat.Docx_2013);
// 关闭文档
document.dispose();