Spire.Office for Java 9.1.4已发布。在该版本中,Spire.PDF for Java 提高了绘制水印的效率;Spire.Doc for Java 增加了添加图像水印的新方法;Spire.Presentation for Java 提升了从PowerPoint 到 SVG格式的转换速度。此外,一些已知问题也在该版本中得到修复。详情请阅读以下内容。
获取 Spire.Office for Java 9.1.4请点击:https://www.e-iceblue.cn/Downloads/Spire-Office-JAVA.html
新功能:
PdfDocument pdf = new PdfDocument();
pdf.loadFromFile("sample.pdf");
PdfPageBase page = pdf.getPages().get(0);
PdfTextReplacer replacer = new PdfTextReplacer(page);
PdfTextReplaceOptions options= new PdfTextReplaceOptions();
options.setReplaceType(EnumSet.of(ReplaceActionType.WholeWord));
replacer.replaceText("www.google.com", "1234567");
pdf.saveToFile(outputFile);PdfImageHelper imageHelper = new PdfImageHelper();
PdfImageInfo[] imageInfoCollection= imageHelper.getImagesInfo(page);
Delete image:
imageHelper.deleteImage(imageInfoCollection[0]);
Extract image:
int index = 0;
for (com.spire.pdf.utilities.PdfImageInfo img : imageInfoCollection) {
BufferedImage image = img.getImage();
File output = new File(outputFile_Img + String.format("img_%d.png", index));
ImageIO.write(image, "PNG", output);
index++;
}PdfImage image = PdfImage.fromFile("ImgFiles/E-iceblue logo.png");
imageHelper.replaceImage(imageInfoCollection[i], image);
Compress image:
for (PdfPageBase page : (Iterable<PdfPageBase>)doc.getPages())
{
if (page != null)
{
if (imageHelper.getImagesInfo(page) != null)
{
for (com.spire.pdf.utilities.PdfImageInfo info : imageHelper.getImagesInfo(page))
{
info.tryCompressImage();
}
}
}
}问题修复:
问题修复:
调整:
新功能:
Document构造中设置newEngine不再起作用,内部默认采用新引擎
HeaderType枚举
GroupedShapeCollection类
ShapeObjectTextCollection类
MailMergeData接口
EnumInterface接口
public PictureWaterMark(InputStream inputeStream,boolean washout)
public PictureWaterMark(String filename,boolean washout)
Field类中downloadImage方法
IDocOleObject接口
PointsConverter类com.spire.license.LicenseProvider -> com.spire.doc.License.LicenseProvider// 设置自定义字体
Document.setCustomFontsFolders(string filePath);
// 处理自定义字体
Document.clearCustomFontsFolders();
// 清除缓存中占用内存的系统字体缓存
Document.clearSystemFontCache();
Example code:
Document doc = new Document();
doc.loadFromFile("inputFile.docx");
doc.setCustomFontsFolders(@"d:\Fonts");
doc.saveToFile("output.pdf", FileFormat.PDF);
doc.close();
doc.dispose();com.spire.doc.FileFormat.WPS -> com.spire.doc.FileFormat.Wps
com.spire.doc.FileFormat.WPT -> com.spire.doc.FileFormat.Wpt
ComparisonLevel -> TextDiffModeComparisonLevel getLevel() -> getTextCompareLevel()
setLevel(ComparisonLevel value) -> setTextCompareLevel(TextDiffMode)
IsPasswordProtect() -> isEncrypted()
getFillEfects() -> getFillEffects()File imageFile = new File("data/E-iceblue.png");
BufferedImage bufferedImage = ImageIO.read(imageFile);
// 通过输入 BufferedImage 创建 PictureWatermark 类的新实例,并设置水印图像的缩放因子
PictureWatermark picture = new PictureWatermark(bufferedImage,false);
// 或者使用另一种创建 PictureWatermark 的方法
// PictureWatermark picture = new PictureWatermark();
// picture.setPicture(bufferedImage);
// picture.isWashout(false);
// 设置水印图片的缩放比例
picture.setScaling(250);
// 设置要应用于文档的水印
document.setWatermark(picture);// 创建一个新的 Document 实例
Document document = new Document();
// 向文档添加一个节
Section section = document.addSection();
// 将段落添加到该节并向其附加文本
section.addParagraph().appendText("Line chart.");
// 添加一个新段落到该节
Paragraph newPara = section.addParagraph();
// 将折线图形状附加到指定宽度和高度的段落
ShapeObject shape = newPara.appendChart(ChartType.Line, 500, 300);
// 从形状中获取图表对象
Chart chart = shape.getChart();
// 获取图表的标题
ChartTitle title = chart.getTitle();
// 设置图表标题的文本
title.setText("My Chart");
// 清除图表中任何现有的系列
ChartSeriesCollection seriesColl = chart.getSeries();
seriesColl.clear();
// 定义类别(X 轴值)
String[] categories = { "C1", "C2", "C3", "C4", "C5", "C6" };
// 将两个具有指定类别和 Y 轴值的系列添加到图表中
seriesColl.add("AW Series 1", categories, new double[] { 1, 2, 2.5, 4, 5, 6 });
seriesColl.add("AW Series 2", categories, new double[] { 2, 3, 3.5, 6, 6.5, 7 });
// 将文档保存为Docx格式的文件
document.saveToFile("AppendLineChart.docx", FileFormat.Docx_2016);
// 释放文档资源
document.dispose();// 创建一个新的 Document 实例
Document doc = new Document();
// 从指定文件加载文档
doc.loadFromFile(inputFile);
// 使用加载的文档创建一个FixedLayoutDocument对象
FixedLayoutDocument layoutDoc = new FixedLayoutDocument(doc);
// 创建一个StringBuilder来存储提取的文本
StringBuilder stringBuilder = new StringBuilder();
// 获取第一页的第一行并将其附加到 StringBuilder
FixedLayoutLine line = layoutDoc.getPages().get(0).getColumns().get(0).getLines().get(0);
stringBuilder.append("Line: " + line.getText() + "\r\n");
// 检索与该行关联的原始段落并将其文本附加到 StringBuilder
Paragraph para = line.getParagraph();
stringBuilder.append("Paragraph text: " + para.getText() + "\r\n");
// 检索第一页上的所有文本,包括页眉和页脚,并将其附加到 StringBuilder
String pageText = layoutDoc.getPages().get(0).getText();
stringBuilder.append(pageText + "\r\n");
// 遍历文档中的每一页并打印每页的行数
for (Object obj : layoutDoc.getPages()) {
FixedLayoutPage page = (FixedLayoutPage) obj;
LayoutCollection<LayoutElement> lines = page.getChildEntities(LayoutElementType.Line, true);
stringBuilder.append("Page " + page.getPageIndex() + " has " + lines.getCount() + " lines." + "\r\n");
}
// 对第一段的布局实体执行反向查找并将它们附加到 StringBuilder
stringBuilder.append("\r\n");
stringBuilder.append("The lines of the first paragraph:" + "\r\n");
for (Object object : layoutDoc.getLayoutEntitiesOfNode(((Section) doc.getFirstChild()).getBody().getParagraphs().get(0))) {
FixedLayoutLine paragraphLine = (FixedLayoutLine) object;
stringBuilder.append(paragraphLine.getText().trim() + "\r\n");
stringBuilder.append(paragraphLine.getRectangle().toString() + "\r\n");
stringBuilder.append("");
}
// 将提取的文本写入文件
FileWriter fileWriter = new FileWriter(new File(outputFile));
fileWriter.write(stringBuilder.toString());
fileWriter.flush();
fileWriter.close();
// 释放文档资源
doc.close();
doc.dispose();// 创建一个新的文档对象
Document document = new Document();
// 在文档中添加一个新的节
Section section = document.addSection();
// 添加一个新的段落到该节
Paragraph paragraph = section.addParagraph();
// 将图片 (SVG) 附加到段落中
paragraph.appendPicture(inputSvg);
// 将文档保存到指定的输出文件
document.saveToFile(outputFile, FileFormat.Docx_2013);
// 关闭文档
document.dispose();问题修复:
问题修复:
新功能:
presentation.loadFromStream(inputStream, FileFormat.AUTO,"password"); Presentation ppt = new Presentation();
ISlide slide = ppt.getSlides().get(0);
List<Point2D> points = new ArrayList<>();
points.add(new Point2D.Float(50f, 50f));
points.add(new Point2D.Float(50f, 150f));
points.add(new Point2D.Float(60f, 200f));
points.add(new Point2D.Float(200f, 200f));
points.add(new Point2D.Float(220f, 150f));
points.add(new Point2D.Float(150f, 90f));
points.add(new Point2D.Float(50f, 50f));
IAutoShape autoShape = slide.getShapes().appendFreeformShape(points);
autoShape.getFill().setFillType(FillFormatType.NONE);
ppt.saveToFile("out.pptx", FileFormat.PPTX_2013);
ppt.dispose();Presentation ppt = new Presentation();
ppt.getSlides().get(0).getShapes().appendShape(ShapeType.LINE, new Point2D.Float(50, 70), new Point2D.Float(150, 120));
ppt.saveToFile( "result.pptx ,FileFormat.PPIX_2013),
ppt.dispose().
在数据分析、财务统计、项目管理等多种工作场景中,经常会遇到需要将多个Excel文件或工作表合并的需求。手动逐一复制粘贴不仅耗时费力,还容易出现遗漏和错误。借助编程语言实现Excel文件的批量合并,能够极大提升工作效率和数据准确性。
本文将详细介绍如何使用 Python 和 Spire.XLS for Python 库来实现多Excel文件和多工作表的批量合并。无论你是数据分析师、财务人员,还是项目经理,这篇教程都能为你带来实用的解决方案,让复杂的 Excel 合并工作变得轻松高效。
使用 Python 合并 Excel 文件有以下几个优势:
Spire.XLS for Python 是一款独立且功能强大的 Excel 处理库,专为创建、读取、编辑和转换 Excel 文件而设计。它无需依赖 Microsoft Excel 软件即可完成各类操作,非常适合服务器或无 Excel 环境下的自动化任务。
主要功能包括:
安装方法
在终端或命令行中执行以下命令即可快速安装 Spire.XLS:
pip install spire.xls
将多个 Excel 文件合并到一个文件中,可以方便统一管理和查看数据,特别适用于整合来自不同部门、区域或时间段的工作表。该方法会完整保留每个原始文件中的所有工作表,确保数据结构和内容不受影响。
实现步骤
import os
from spire.xls import *
# 定义待合并Excel文件所在的文件夹路径
input_folder = './excel_files'
# 定义合并后保存的Excel文件名称
output_file = 'merged_workbook.xlsx'
# 创建目标工作簿,用于保存合并后的所有工作表
merged_workbook = Workbook()
# 删除默认工作表
merged_workbook.Worksheets.Clear()
# 遍历文件夹中所有文件
for filename in os.listdir(input_folder):
# 只处理扩展名为xls或xlsx的文件(忽略大小写)
if filename.lower().endswith(('.xls', '.xlsx')):
file_path = os.path.join(input_folder, filename)
# 加载当前Excel文件
source_wb = Workbook()
source_wb.LoadFromFile(file_path)
# 复制当前文件中的每个工作表到目标工作簿
for i in range(source_wb.Worksheets.Count):
sheet = source_wb.Worksheets[i]
merged_workbook.Worksheets.AddCopy(sheet, WorksheetCopyType.CopyAll)
# 释放当前工作簿资源
source_wb.Dispose()
# 保存合并后的工作簿到指定文件
merged_workbook.SaveToFile(output_file, ExcelVersion.Version2016)
print(f"合并完成,文件已保存为 {output_file}")
# 释放目标工作簿资源
merged_workbook.Dispose()

将多个 Excel 工作表的数据汇总到同一个工作表中,有助于集中管理和分析销售记录、调查数据、绩效报表等信息,提高数据整合效率。
import os
from spire.xls import *
# 定义存放Excel文件的文件夹路径
input_folder = './excel_worksheets'
# 定义合并后保存的文件名
output_file = 'merged_into_one_sheet.xlsx'
# 创建一个新的工作簿,并移除默认工作表
merged_workbook = Workbook()
merged_workbook.Worksheets.Clear()
# 创建一个新的工作表,作为目标工作表
merged_sheet = merged_workbook.Worksheets.Add("Sheet1")
current_row = 1 # 目标工作表当前写入的起始行
# 遍历文件夹中的所有Excel文件
for filename in os.listdir(input_folder):
if filename.lower().endswith(('.xlsx', '.xls')):
file_path = os.path.join(input_folder, filename)
# 加载当前Excel文件
workbook = Workbook()
workbook.LoadFromFile(file_path)
# 获取当前文件的第一个工作表
sheet = workbook.Worksheets[0]
# 获取当前工作表的有效数据范围(只包含有数据的区域)
source_range = sheet.AllocatedRange
# 目标工作表的起始写入单元格
dest_range = merged_sheet.Range[current_row, 1]
# 将数据复制到目标工作表指定位置
source_range.Copy(dest_range)
# 更新当前写入行数(按实际数据行数累加)
current_row += source_range.RowCount
# 释放当前工作簿资源
workbook.Dispose()
# 保存合并后的工作簿
merged_workbook.SaveToFile(output_file, ExcelVersion.Version2016)
merged_workbook.Dispose()
print(f"多个工作表数据已合并至一个工作表,文件保存为 {output_file}")

通过 Python 和 Spire.XLS 库,可以实现对 Excel 文件的自动化批量合并,既节省了大量的人工操作时间,也降低了因手动处理带来的错误风险。该方法支持多种 Excel 格式,且不依赖于 Microsoft Excel 软件,适合多种运行环境。无论是将多个文件的工作表合并到一个文件,还是将多个工作表的数据汇总到单一工作表,自动化处理都能有效提升数据管理的效率和准确性。
A1:可以,Spire.XLS 支持这两种格式,无需转换。
A2:不需要,Spire.XLS 是独立库,无需安装 Microsoft Office。
A3:可以,代码中可按工作表名称或索引选择,例如:
sheet = source_workbook.Worksheets["Summary"]
A4:可以添加判断逻辑,例如:
if current_row > 1:
start_row = 2 # 跳过表头
else:
start_row = 1
A5:可以,在合并表中新增一列,用于记录来源文件名。
A6:其限制与 Excel 本身一致,.xlsx 支持最大 1,048,576 行 × 16,384 列,.xls 支持最大 65,536 行 × 256 列。
A7:可以,合并过程中会完整保留公式和格式。
Spire.Doc for Java 12.1.10已发布。本次更新将命名空间com.spire.ms.Printing.*更改为com.spire.doc.printing.*。此外,一些已知问题也在该版本中被成功修复,如将Word转换为PDF时,程序抛出java.lang.OutOfMemoryError异常的问题。详情请阅读以下内容。
功能调整:
问题修复:
Spire.XLS for Java 14.1.1 已发布。本次更新增强了 HTML 到 XLSX 的转换功能。此外,一些已知问题也在该版本中得到修复,如用WPS工具进行打印预览被保存出的XLSX文档时页边距不正确的问题。详情请阅读以下内容。
问题修复:
Spire.Doc 12.1.5 已发布。该版本移除了 MonoAndroid 包和 Xamarin.iOS 包。同时也对文档的密码功能做出了一系列变更。此外,还修复了一些已知问题,如修复了在新建 Document() 对象和应用 LicenseKey 后,加载文档时部分数据丢失的问题。详情请阅读以下内容。
功能调整:
问题修复:
Spire.Office for Python 9.1.0 已发布。本次更新在 Spire.PDF for Python、Spire.Doc for Python、Spire.XLS for Python 和 Spire.Presentation for Python 中增加了自定义异常类 SpireException,同时还修复了一些已知问题。详情请阅读以下内容。
获取 Spire.Office for Python 9.1.0请点击:https://www.e-iceblue.cn/Downloads/Spire-Office-Python.html
新功能:
问题修复:
新功能:
问题修复:
新功能:
新功能:
表格作为一种强大的数据可视化工具,在 PowerPoint 演示文稿中扮演着关键角色。其设计基于行和列的有序结构,使得各类文本、数值和其他类型的信息能够在单元格内精确呈现。在制作演示文稿时,插入表格能够让用户有效地构建和展示经过整理的数据内容,从而使幻灯片显得更有逻辑层次。相较于单纯的文本叙述,表格更能凸显各组数据之间的对比关系,增强数据的易读性与直观性,进而提升观众对演示内容的理解深度和接受度。本文将指导如何运用 Spire.Presentation for Python 在 PowerPoint 中添加和编辑表格。
本教程需要 Spire.Presentation for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.Presentation如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.Presentation for Python
Spire.Presentation for Python 提供了 Presentation.Slides[].Shapes.AppendTable(x: float, y: float, widths: List[float], heights: List[float]) 方法,用于向 PowerPoint 演示文稿中添加表格。具体操作步骤如下:
from spire.presentation.common import *
import math
from spire.presentation import *
inputFile = "模板.pptx"
outputFile = "创建表格.pptx"
# 创建 Presentation 类的对象
presentation = Presentation()
# 加载一个演示文稿
presentation.LoadFromFile(inputFile)
# 定义表格列宽度
widths = [100, 100, 100, 100, 100]
# 定义表格行高度
heights = [25, 25, 25, 25, 25, 25, 25, 25, 25]
# 计算表格左边距离
left = math.trunc(presentation.SlideSize.Size.Width / float(2)) - 250
# 在第一页幻灯片上添加表格
table = presentation.Slides[0].Shapes.AppendTable(left, 150, widths, heights)
# 将表格数据定义为二维字符串数组
dataStr = [["产品ID", "产品名称", "销售数量(斤)","单价(元/斤)","销售额(元)"],
["001", "香蕉", "800","3.0","2400.00"],
["002", "苹果", "500","6.5","3250.00"],
["003", "菠萝", "500","5.0","2500.00"],
["004", "芒果", "200","6.0","1200.00"],
["005", "橙子", "1000","5.0","5000.00"],
["006", "柠檬", "300","7.0","2100.00"],
["007", "枇杷", "120","10.0","1200.00"],
["008", "草莓", "600","30.00","18000.00"]]
# 循环遍历数组
for i in range(0, len(dataStr)):
for j in range(0, len(dataStr[0])):
# 使用这些数据填充表格的每个单元格
table[j,i].TextFrame.Text = dataStr[i][j]
# 设置字体名称和字体大小
table[j,i].TextFrame.Paragraphs[0].TextRanges[0].LatinFont = TextFont("微软雅黑")
table[j,i].TextFrame.Paragraphs[0].TextRanges[0].FontHeight = 12
# 将表格的第一行对齐方式设置为居中
for i in range(0, len(dataStr[0])):
table[i,0].TextFrame.Paragraphs[0].Alignment = TextAlignmentType.Center
# 应用样式到表格
table.StylePreset = TableStylePreset.LightStyle3Accent1
# 保存结果文件
presentation.SaveToFile(outputFile, FileFormat.Pptx2013)
# 释放对象
presentation.Dispose()
您还可以根据需要编辑演示文稿中的表格,例如替换数据、更改样式、突出显示数据等。以下是详细步骤:
from spire.presentation.common import *
from spire.presentation import *
inputFile = "创建表格.pptx"
outputFile = "编辑表格.pptx"
# 创建Presentation类的对象
presentation = Presentation()
# 加载示例演示文稿
presentation.LoadFromFile(inputFile)
# 需要替换的数据数组
strs = ["002", "水蜜桃", "300", "8.0", "2400.00"]
table = None
# 遍历第一张幻灯片中的形状
for shape in presentation.Slides[0].Shapes:
# 判断形状是否为表格
if isinstance(shape, ITable):
table = shape
# 设置表格样式
table.StylePreset = TableStylePreset.LightStyle1Accent2
# 使用循环替换特定单元格的数据
for i, unusedItem in enumerate(table.ColumnsList):
# 替换单元格中的数据
table[i,2].TextFrame.Text = strs[i]
# 高亮显示新数据
table[i,2].TextFrame.TextRange.HighlightColor.Color = Color.get_Yellow()
# 保存结果文件
presentation.SaveToFile(outputFile, FileFormat.Pptx2013)
# 释放对象
presentation.Dispose()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
在创建新 PDF 文档时,通过页面布局设计,我们可以在 PDF 页面上下空白处添加一些公司信息,图标,页码等信息作为页眉页脚,以增强 PDF 文档的外观效果和专业性。本文将介绍如何使用 Spire.PDF for Python 创建新 PDF 时添加页眉页脚。
本教程需要 Spire.PDF for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.PDF如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.PDF for Python
Spire.PDF for Python 提供 PdfPageTemplateElement 类,用于定义页面模板元素,并通过提供的方法 PdfPageTemplateElement.Graphics.DrawString(),PdfPageTemplateElement.Graphics.DrawLine(),PdfPageTemplateElement.Graphics.DrawImage() 等绘制相应的文本,线条,图片内容,同时还支持用 PdfGraphicsWidget.Draw() 绘制如 PdfPageCountField、PdfPageNumberField 等动态字段域到此模板元素。
PdfPageTemplateElement 模板元素上绘制内容,坐标体系设定如下:
Spire.PDF for Python 提供 PdfDocumentTemplate 类,用于 PDF 整个页面模板设计,上面定义的 PdfPageTemplateElement 页面模板元素可以直接应用到 PdfDocumentTemplate 页面模板上,PdfDocumentTemplate 页面模板可以应用1个或多个 PdfPageTemplateElement 页面模板元素,比如应用到 PdfDocumentTemplate.Top 和 PdfDocumentTemplate.Bottom 页面模板的上下区域实现 PDF 页眉,页脚的效果。
Spire.PDF 新建的 PDF 页面默认含有边距,PdfDocumentTemplate 页面模板初始化坐标设定如下:

页边距的地方不能绘制内容,要应用 PdfPageTemplateElement 到 PdfDocumentTemplate 实现页眉页脚效果,可以先重置 PDF 页面边距为0,这样新建 PDF 页面 PdfDocumentTemplate 页面模板坐标体系会依据 PdfPageTemplateElement 设置的大小调整,例如:

以下是使用 Spire.PDF for Python 创建新 PDF 时实现在页眉添加文本、图片以及线条:
第一部分:自定义方法 CreateHeaderTemplate() 设计页眉模板元素
第二部分:创建 PDF 文档对象并调用上面自定义方法添加页眉
from spire.pdf.common import *
from spire.pdf import *
#定义CreateHeaderTemplate()方法
def CreateHeaderTemplate(doc, pageSize, margins):
#创建指定大小的页眉模板
headerSpace = PdfPageTemplateElement(pageSize.Width, margins.Top)
headerSpace.Foreground = True
doc.Template.Top = headerSpace
#初始化x,y 坐标点
x = margins.Left
y = 0.0
# 设置字体,画刷,画笔,文本对齐格式
font = PdfTrueTypeFont("宋体", 10.0, PdfFontStyle.Italic, True)
brush = PdfBrushes.get_Gray()
pen = PdfPen(PdfBrushes.get_Gray(), 1.0)
leftAlign = PdfTextAlignment.Left
# 加载页眉图片并获取图片point高宽值
headerImage = PdfImage.FromFile("header.png")
width = headerImage.Width
height = headerImage.Height
unitCvtr = PdfUnitConvertor()
pointWidth = unitCvtr.ConvertUnits(width, PdfGraphicsUnit.Pixel, PdfGraphicsUnit.Point)
pointHeight = unitCvtr.ConvertUnits(height, PdfGraphicsUnit.Pixel, PdfGraphicsUnit.Point)
# 在指定位置绘制页眉图片
headerSpace.Graphics.DrawImage(headerImage, headerSpace.Width-x-pointWidth, headerSpace.Height-pointFeight)
# 在指定位置绘制页眉文本
headerSpace.Graphics.DrawString("成都冰蓝科技有限公司\nwww.e-iceblue.cn", font, brush, x, headerSpace.Height-font.Height*2, PdfStringFormat(leftAlign))
# 在指定位置绘制页眉线条
headerSpace.Graphics.DrawLine(pen, x, margins.Top, pageSize.Width - x, margins.Top)
# 创建 PdfDocument 对象
doc = PdfDocument()
# 设置页面大小和边距
pageSize =PdfPageSize.A4()
doc.PageSettings.Size = pageSize
doc.PageSettings.Margins = PdfMargins(0.0)
# 新建PdfMargins对象用于设置页眉,页脚及左右模板大小
margins = PdfMargins(50.0, 50.0, 50.0, 50.0)
doc.Template.Left = PdfPageTemplateElement(margins.Left, pageSize.Height-margins.Bottom-margins.Top)
doc.Template.Right = PdfPageTemplateElement(margins.Right, pageSize.Height-margins.Bottom-margins.Top)
doc.Template.Bottom = PdfPageTemplateElement(pageSize.Width, margins.Bottom)
# 调用 CreateHeaderTemplate()添加页眉
CreateHeaderTemplate(doc, pageSize, margins)
# 按以上设置添加页面
page = doc.Pages.Add()
# 定义页面内容要使用的字体,画刷
font = PdfTrueTypeFont("宋体", 14.0, PdfFontStyle.Regular, True)
brush = PdfBrushes.get_Blue()
# 绘制文本内容到页面
text = "Spire.PDF 添加页眉示例"
page.Canvas.DrawString(text, font, brush, 0.0, 20.0)
# 将文档保存为 PDF 格式
doc.SaveToFile("结果.pdf", FileFormat.PDF)
# 释放文档对象
doc.Close()
以下是使用 Spire.PDF for Python 创建新 PDF 时实现在页脚添加文本、线条以及页码内容:
第一部分:自定义方法 CreateFooterTemplate() 设计页脚模板元素
第二部分:创建 PDF 文档对象并调用上面自定义方法添加页脚
from spire.pdf.common import *
from spire.pdf import *
#定义CreateFooterTemplate()方法
def CreateFooterTemplate(doc, pageSize, margins):
# 创建指定大小的页脚模板
footerSpace = PdfPageTemplateElement(pageSize.Width, margins.Bottom)
footerSpace.Foreground = True
doc.Template.Bottom = footerSpace
#初始化x,y 坐标点
x = margins.Left
y = 0.0
# 设置字体,画刷,画笔,文本对齐格式
font = PdfTrueTypeFont("宋体", 12.0, PdfFontStyle.Italic, True)
brush = PdfBrushes.get_Gray()
pen = PdfPen(PdfBrushes.get_Gray(), 1.0)
leftAlign = PdfTextAlignment.Left
# 在指定位置绘制页脚线条
footerSpace.Graphics.DrawLine(pen, x, y, pageSize.Width - x, y)
# 在指定位置绘制页脚文本
footerSpace.Graphics.DrawString("邮箱:sales @e-iceblue.com\n电话:028-81705109 ", font, brush, x, y, PdfStringFormat(leftAlign))
# 创建页码编号和总页数字段域
number = PdfPageNumberField()
count = PdfPageCountField()
listAutomaticField = [number, count]
# 创建复合字段域并设置字符串格式进行绘制
compositeField = PdfCompositeField(font, PdfBrushes.get_Gray(), "第{0}页共{1}页", listAutomaticField)
compositeField.StringFormat = PdfStringFormat(PdfTextAlignment.Right, PdfVerticalAlignment.Top)
size = font.MeasureString(compositeField.Text)
compositeField.Bounds = RectangleF(pageSize.Width -x-size.Width, y, size.Width, size.Height)
newTemplate = compositeField
templateGraphicsWidget = PdfGraphicsWidget(newTemplate.Ptr)
templateGraphicsWidget.Draw(footerSpace.Graphics)
# 创建 PdfDocument 对象
doc = PdfDocument()
# 设置页面大小和边距
pageSize =PdfPageSize.A4()
doc.PageSettings.Size = pageSize
doc.PageSettings.Margins = PdfMargins(0.0)
# 新建PdfMargins对象用于设置页眉,页脚及左右模板大小
margins = PdfMargins(50.0, 50.0, 50.0, 50.0)
doc.Template.Left = PdfPageTemplateElement(margins.Left, pageSize.Height-margins.Top-margins.Bottom)
doc.Template.Right = PdfPageTemplateElement(margins.Right, pageSize.Height-margins.Top-margins.Bottom)
doc.Template.Top = PdfPageTemplateElement(pageSize.Width, margins.Top)
# 调用 CreateFooterTemplate()添加页脚
CreateFooterTemplate(doc, pageSize, margins)
# 按以上设置添加页面
page = doc.Pages.Add()
# 创建页面内容要使用的字体,画刷
font = PdfTrueTypeFont("宋体", 14.0, PdfFontStyle.Regular, True)
brush = PdfBrushes.get_Blue()
# 绘制文本内容到页面
text = "Spire.PDF 添加页脚示例"
page.Canvas.DrawString(text, font, brush, 0.0, pageSize.Height-margins.Bottom-margins.Top-font.Height-20)
# 将文档保存为 PDF 格式
doc.SaveToFile("结果.pdf", FileFormat.PDF)
# 释放文档对象
doc.Close()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。

将 PDF 内容转换为 HTML,不仅能够让文档在网页上轻松访问,还能显著提升可用性、搜索性和跨设备兼容性。无论您是在开发 PDF 查看器、自动化文档工作流,还是进行内容的在线发布,使用 Python 将 PDF 转换为 HTML 都能有效提升用户体验。
本教程将详细介绍如何使用 Python 将 PDF 转换为 HTML,从基础的转换操作到进阶的自定义设置,再到基于流的输出方式。每个部分都附有实用的代码示例,帮助您快速理解和完成 PDF 到 HTML 的转换。
HTML(超文本标记语言)是网页内容的基础语言。将 PDF 转换为 HTML,能够让文档内容在网页上更加方便地浏览、编辑和索引。将 PDF 导出为 HTML 的主要优点包括:
在将 PDF 转换为 HTML之前,您需要安装支持处理 PDF 文档并导出为HTML 格式的库。在本教程中,我们将使用 Spire.PDF for Python,它是一个高性能的 PDF 库,支持多种PDF 文档处理和转换功能,并且不依赖第三方软件。
安装Spire.PDF for Python
您可以通过 pip 安装 Spire.PDF for Python,只需在终端中执行以下命令:
pip install Spire.PDF
该命令将自动下载并安装最新版本的 Spire.PDF 包及其依赖项。
如果您需要安装帮助,可以参考这篇教程:如何在 Windows 中安装 Spire.PDF for Python。
Spire.PDF 提供了 SaveToFile() 方法,可以轻松地将整个 PDF 文档快速导出为 HTML 格式。此方法能够保留 PDF 文档的原始布局和结构,使得转换后的 HTML 文件在网页上呈现出与原始 PDF 一样的效果。
以下是一个基本的 PDF 转 HTML 的代码示例:
from spire.pdf.common import *
from spire.pdf import *
# 初始化 PdfDocument 对象
doc = PdfDocument()
# 加载 PDF 文件
doc.LoadFromFile("示例.pdf")
# 将 PDF 转换并保存为 HTML
doc.SaveToFile("output/Pdf转Html.html", FileFormat.HTML)
# 关闭文档
doc.Close()
下图展示了转换前的 PDF 文件和生成后的 HTML 文件效果:

如果您希望在转换过程中对 HTML 输出进行更精细的控制,可以使用 SetPdfToHtmlOptions() 方法进行设置。该方法提供了多个参数,允许您定制转换效果,包括图像嵌入、每个文件输出的页面数量以及 SVG 图像的质量等。
以下是主要参数及其功能:
| 参数 | 类型 | 描述 |
|---|---|---|
| useEmbeddedSvg | bool | 如果为 True,则嵌入 SVG 图像 |
| useEmbeddedImg | bool | 如果为 True,则嵌入图片(仅在 useEmbeddedSvg 设置为 False 时生效) |
| maxPageOneFile | bool | 限制每个 HTML 文件仅输出一页内容(仅在 useEmbeddedSvg 设置为 False 时生效) |
| useHighQualityEmbeddedSvg | bool | 启用高分辨率的 SVG 图像(仅在 useEmbeddedSvg 设置为 True 时生效) |
代码示例:
from spire.pdf.common import *
from spire.pdf import *
# 初始化 PdfDocument 对象
doc = PdfDocument()
# 加载 PDF 文件
doc.LoadFromFile("示例.pdf")
# 获取转换设置
options = doc.ConvertOptions
# 自定义转换:使用图像嵌入,每个文件一页
options.SetPdfToHtmlOptions(False, True, 1, False)
# 保存 PDF 为自定义选项的 HTML 文件
doc.SaveToFile("output/PDF转HTML设置选项.html", FileFormat.HTML)
# 关闭文档
doc.Close()
在 Web 或云应用中,您可能更希望将 HTML 输出写入流(例如通过 HTTP 提供服务),而不是直接保存到文件系统。此时,您可以使用 SaveToStream() 方法来实现这一需求。
代码示例:
from spire.pdf.common import *
from spire.pdf import *
# 初始化 PdfDocument 对象
doc = PdfDocument()
# 加载 PDF 文件
doc.LoadFromFile("示例.pdf")
# 创建流来保存 HTML 输出
fileStream = Stream("output/PDF转HTML流.html")
# 将 PDF 保存为 HTML 流
doc.SaveToStream(fileStream, FileFormat.HTML)
# 关闭流和文档
fileStream.Close()
doc.Close()
使用 Python 将 PDF 转换为 HTML 是让文档更好地适配网页并提升互动性的理想方式。通过 Spire.PDF for Python,您可以全面掌控转换过程,无论是简单导出,还是嵌入图像、SVG,甚至流式输出等高级选项,都可以轻松实现。
A1: 使用 Spire.PDF,您可以通过 doc.LoadFromFile("file.pdf", "password") 打开受密码保护的 PDF,并将其成功转换为 HTML 格式。
A2: 支持。默认情况下,Spire.PDF 会将 PDF 文件中的所有页面转换为 HTML。您还可以通过 maxPageOneFile 参数设置每个 HTML 文件显示多少页,以满足不同需求。
A3: 会的,Spire.PDF 会根据您的转换设置(如图像或 SVG 嵌入)尽可能保留图像和字体,确保 HTML 输出与原 PDF 的视觉效果一致。
如果您希望在没有评估限制的情况下全面体验 Spire.PDF for Python 的功能,可以申请免费的 30 天试用许可证。
在 PowerPoint 中对形状进行分组有助于简化复杂的编辑任务,尤其适用于批量修改格式或定位。分组后可一次性调整整体,大大提升效率;而取消分组则能恢复对单个形状的独立操控,便于进行精细化编辑与个性化设置。本文将指导如何运用 Spire.Presentation for Python 在 PowerPoint 中实现形状的分组与取消分组。
本教程需要 Spire.Presentation for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.Presentation如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.Presentation for Python
Spire.Presentation for Python 提供了 ISlide.GroupShapes(shapeList: List) 方法,用于在特定幻灯片上组合两个或多个形状。具体操作步骤如下:
from spire.presentation import *
# 创建一个Presentation对象
ppt = Presentation()
# 获取演示文稿中的第一张幻灯片
slide = ppt.Slides[0]
# 添加一个填充为天蓝色的矩形
rectangle = slide.Shapes.AppendShape(ShapeType.Rectangle, RectangleF.FromLTRB (250, 180, 450, 220))
rectangle.Fill.FillType = FillFormatType.Solid
rectangle.Fill.SolidColor.KnownColor = KnownColors.SkyBlue
# 设置矩形边框线宽为0.1单位
rectangle.Line.Width = 0.1
# 添加一个填充为浅粉色的丝带形状
ribbon = slide.Shapes.AppendShape(ShapeType.Ribbon2, RectangleF.FromLTRB (290, 155, 410, 235))
ribbon.Fill.FillType = FillFormatType.Solid
ribbon.Fill.SolidColor.KnownColor = KnownColors.LightPink
# 设置矩形边框线宽为0.1单位
ribbon.Line.Width = 0.1
# 将这两个形状添加到列表中
shape_list = []
shape_list.append(rectangle)
shape_list.append(ribbon)
# 组合列表中的形状
slide.GroupShapes(shape_list)
# 保存文档
ppt.SaveToFile("组合.pptx", FileFormat.Pptx2013)
# 释放对象
ppt.Dispose()
要取消 PowerPoint 文档中的组合形状,您可以遍历文档中的所有幻灯片以及每张幻灯片上的所有形状,找出已分组的形状并使用 ISlide.Ungroup(groupShape: GroupShape) 方法将其取消组合。具体的步骤如下:
from spire.presentation import *
# 创建一个Presentation类的对象
ppt = Presentation()
# 加载一个包含组合形状的PowerPoint文档
ppt.LoadFromFile("组合.pptx")
# 遍历文档中的所有幻灯片
for i in range(ppt.Slides.Count):
slide = ppt.Slides[i]
# 遍历该幻灯片上的所有形状
for j in range(slide.Shapes.Count):
shape = slide.Shapes[j]
# 判断是否为组合形状
if isinstance(shape, GroupShape):
groupShape = shape
# 取消组合该组合形状
slide.Ungroup(groupShape)
# 保存文档
ppt.SaveToFile("取消组合.pptx", FileFormat.Pptx2013)
# 释放对象
ppt.Dispose()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。