Spire.Presentation for Java 10.12.4 已发布。该版本主要修复了一个公式在 Mac 上显示不正确的问题。详情请查看以下内容。
问题修复:
https://www.e-iceblue.cn/Downloads/Spire-Presentation-JAVA.html

在 Java 应用中,PDF 解析(PDF parsing in Java)通常用于从 PDF 文件中提取可用信息,而不仅仅是将其渲染出来进行展示。常见的应用场景包括文档索引、自动化报表处理、发票分析以及数据采集与导入流程等。
与 JSON、XML 等结构化数据格式不同,PDF 的设计目标是保证视觉呈现效果的一致性。文本、表格、图像等内容在 PDF 中并不是以逻辑结构存储的,而是以带有坐标信息的绘制指令形式存在。因此,在 Java 中进行 PDF 解析,核心在于理解 PDF 内部的内容表示方式,以及 Java PDF 库是如何通过 API 将这些内容暴露出来的。
本文将基于 Spire.PDF for Java,从实际开发角度出发,介绍在 Java 项目中常见的 PDF 解析操作。文章不会将 PDF 解析视为一个单一的线性流程,而是按功能划分,分别讲解文本、表格、图像和元数据的提取方式,便于在真实项目中按需组合使用。
目录
从实践层面来看,Java 中的 PDF 解析并不是一个单一操作,而是一组针对同一 PDF 文档执行的不同数据提取任务,具体取决于应用需要获取哪类信息。
在实际系统中,PDF 解析通常用于获取以下内容:
PDF 解析之所以复杂,根本原因在于 PDF 的内容存储方式。与结构化文档不同,PDF 并不会显式保存段落、行或表格等逻辑结构,而是主要由以下内容组成:
因此,Java 中的 PDF 解析本质上是基于页面布局信息还原内容语义的过程。这也是为什么在实际项目中,往往需要借助专业的 PDF 解析库:它既能暴露底层页面内容,又提供了文本提取、表格识别等高级功能,从而减少手写解析逻辑的复杂度。
在生产环境中,PDF 解析更适合被设计为一组可独立调用的解析操作,而不是固定顺序的流水线。这种设计方式有助于隔离错误,也能让应用只执行真正需要的解析逻辑。
本文使用 Spire.PDF for Java 作为示例库。它提供了文本提取、表格解析、图像导出和元数据访问等 API,适用于后端服务、批量任务以及文档自动化系统。
你可以从 Spire.PDF for Java 下载页面 下载并手动引入依赖。如果项目使用 Maven,也可以通过以下配置进行安装:
<repositories>
<repository>
<id>com.e-iceblue</id>
<name>e-iceblue</name>
<url>https://repo.e-iceblue.cn/repository/maven-public/</url>
</repository>
</repositories>
<dependencies>
<dependency>
<groupId>e-iceblue</groupId>
<artifactId>spire.pdf</artifactId>
<version>12.9.0</version>
</dependency>
</dependencies>
完成安装后,即可直接使用 Java 代码加载和解析 PDF 文件,无需依赖外部工具。
在执行任何解析操作之前,首先需要加载并验证 PDF 文档。建议将这一步作为独立操作,用于确认文档是否可以被后续解析逻辑安全处理。
import com.spire.pdf.PdfDocument;
public class loadPDF {
public static void main(String[] args) {
// 创建 PdfDocument 实例
PdfDocument pdf = new PdfDocument();
// 加载 PDF 文件
pdf.loadFromFile("sample.pdf");
// 获取页面总数
int pageCount = pdf.getPages().getCount();
System.out.println("总页数: " + pageCount);
}
}
控制台输出示例

从实现角度来看,只要能够成功加载文档并访问页面集合,就已经验证了多个关键条件:
在生产系统中,这一步通常作为入口校验使用,无法加载或页面结构异常的 PDF 可以直接被拦截,避免影响后续流程,有助于在批处理或自动化场景中避免错误级联。
实际开发中,PDF 也可能以字节数组或流的形式传入。关于这类场景,可参考 使用 Java 从字节数组加载 PDF 文档。
文本解析是 Java 中最常见的 PDF 处理需求之一,其核心目标是从 PDF 页面中提取并重组可读文本内容。使用 Spire.PDF for Java 解析 PDF 文本时,文本解析不应简单理解为一次性 API 调用,而应通过 PdfTextExtractor 配合可配置的 PdfTextExtractOptions 来实现,以获得更稳定、可控的解析结果。
将文本解析设计为独立的处理步骤,可以在文档索引、内容分析、全文搜索或数据迁移等场景中灵活复用。
在典型的 Java 实现中,PDF 文本解析通常由以下几个清晰的步骤构成,每一步都能在代码中直接体现:
这种基于页面的解析方式与 PDF 的底层结构高度一致,也为多页文档的处理提供了更好的控制能力。
import com.spire.pdf.PdfDocument;
import com.spire.pdf.texts.PdfTextExtractOptions;
import com.spire.pdf.texts.PdfTextExtractor;
public class extractPdfText {
public static void main(String[] args) {
// 创建并加载 PDF 文档
PdfDocument pdf = new PdfDocument();
pdf.loadFromFile("sample1.pdf");
// 使用 StringBuilder 高效累积解析结果
StringBuilder extractedText = new StringBuilder();
// 配置文本解析选项
PdfTextExtractOptions options = new PdfTextExtractOptions();
// 启用简化解析模式,提高文本可读性
options.setSimpleExtraction(true);
// 遍历 PDF 中的每一页
for (int i = 0; i < pdf.getPages().getCount(); i++) {
// 为当前页面创建文本解析器
PdfTextExtractor extractor =
new PdfTextExtractor(pdf.getPages().get(i));
// 按配置选项解析当前页面文本
String pageText = extractor.extract(options);
// 追加到结果缓冲区
extractedText.append(pageText).append("\n");
}
// 此时 extractedText 已包含完整文本内容,
// 可用于存储、索引或后续处理
System.out.println(extractedText.toString());
}
}
控制台输出示例

PdfTextExtractor 以页面为解析单位,对文本内容进行提取和重组,提供比全局提取更精细的控制能力。
PdfTextExtractOptions 用于控制文本解析行为。启用 setSimpleExtraction(true) 可在多数场景下生成更干净、更易读的文本结果,减少因布局干扰导致的断行或错序问题。
按页解析策略 将解析范围限定在单页,有助于处理大体量 PDF 文档,也更容易在出现异常时定位和隔离问题页面。
这种基于页面的文本解析方式非常适合布局相对稳定、以文字内容为主的文档,如报告、合同和说明文档;也是在 Java 中使用 Spire.PDF for Java 解析 PDF 页面文本的推荐实践。更多文本解析示例可参考:使用 Java 从 PDF 页面中提取文本。
表格解析属于较为高级的 PDF 解析操作,其目标是在 PDF 页面中识别出表格结构,并将其还原为具有行和列关系的结构化数据。与纯文本解析相比,表格解析更强调单元格之间的语义关系,常用于发票、财务报表、业务统计报表等场景。
在 Java 中进行 PDF 解析时,表格解析可以将视觉上对齐的数据转换为程序可直接处理的结构化内容,便于存储、分析或导出。
表格解析的核心不再是简单的文本提取,而是基于页面布局和对齐关系进行结构推断:
与文本解析不同,表格解析是通过元素的视觉对齐和布局一致性来推断结构的,从而实现对原本“散落在页面上的文本”的行列级访问。
下面的示例展示了如何使用 PdfTableExtractor 从 PDF 页面中解析表格,并将其转换为按行列组织的数据结构,便于后续处理或导出。
import com.spire.pdf.PdfDocument;
import com.spire.pdf.utilities.PdfTable;
import com.spire.pdf.utilities.PdfTableExtractor;
public class extractPdfTable {
public static void main(String[] args) {
// 载入 PDF 文档
PdfDocument pdf = new PdfDocument();
pdf.loadFromFile("sample1.pdf");
// 创建 PdfTableExtractor 对象
PdfTableExtractor extractor = new PdfTableExtractor(pdf);
// 从第一页解析表格(页索引从 0 开始)
PdfTable[] tables = extractor.extractTable(0);
// 遍历表格
if (tables != null) {
for (PdfTable table : tables) {
// 获取表格的行数和列数
int rowCount = table.getRowCount();
int columnCount = table.getColumnCount();
System.out.println("Rows: " + rowCount +
", Columns: " + columnCount);
StringBuilder tableData = new StringBuilder();
for (int i = 0; i < rowCount; i++) {
for (int j = 0; j < columnCount; j++) {
// 获取单元格数据
tableData.append(table.getText(i, j));
if (j < columnCount - 1) {
tableData.append("\t");
}
}
if (i < rowCount - 1) {
tableData.append("\n");
}
}
System.out.println(tableData.toString());
}
}
}
}
控制台输出示例

PdfTableExtractor: 通过分析页面级内容,根据文本对齐和布局特征识别表格区域。
结构还原: 行和列是通过文本元素的相对位置推断得出的,可通过行列索引访问单元格内容。
按页解析: 每页单独解析,有助于应对不同页面布局不一致的问题。
尽管存在一定限制,表格解析依然是 Java PDF 解析中极具价值的能力之一,特别适合从结构化业务文档中自动提取数据。
在解析出表格数据后,常见的做法是将其导出为 CSV 等结构化格式,可参考:使用 Java 将 PDF 表格转换为 CSV。
图像解析是一种专门用于提取 PDF 页面中嵌入图像资源的解析能力。与文本或表格解析不同,图像解析并不依赖内容流或布局推断,而是直接分析页面资源,识别其中的图像对象。
在 Java PDF 处理系统中,图像解析常用于视觉内容归档、文档组成审计,或将图像数据传递给下游处理流程。
从实现角度来看,图像解析主要基于页面级资源:
由于图像作为独立资源存储,这一解析过程不依赖文本顺序、布局重建或表格识别逻辑。
import com.spire.pdf.PdfDocument;
import com.spire.pdf.utilities.PdfImageHelper;
import com.spire.pdf.utilities.PdfImageInfo;
import javax.imageio.ImageIO;
import java.awt.image.BufferedImage;
import java.io.File;
import java.io.IOException;
public class extractPdfImages {
public static void main(String[] args) throws IOException {
// 载入 PDF 文档
PdfDocument pdf = new PdfDocument();
pdf.loadFromFile("sample1.pdf");
// 创建 PdfImageHelper 对象
PdfImageHelper imageHelper = new PdfImageHelper();
// 遍历每一页
for (int i = 0; i < pdf.getPages().getCount(); i++) {
// 获取当前页的图片信息
PdfImageInfo[] imageInfos =
imageHelper.getImagesInfo(pdf.getPages().get(i));
if (imageInfos != null) {
for (int j = 0; j < imageInfos.length; j++) {
// 获取指定图片
BufferedImage image = imageInfos[j].getImage();
// 保存图片为 PNG 文件
File output = new File(
"output/images/page_" + i + "_image_" + j + ".png"
);
ImageIO.write(image, "PNG", output);
}
}
}
}
}
提取结果示例

PdfImageHelper / PdfImageInfo: 用于分析页面资源并以 BufferedImage 形式访问嵌入图像。
按页处理: 即使多页文档中存在重复或复用图像,也能准确提取。
与布局无关: 图像解析不依赖文本流或表格结构,适用于所有视觉资源。
除逐个提取图像外,也可以将整个 PDF 页面直接转换为图片,详见:使用 Java 将 PDF 页面转换为图片。
元数据解析是 PDF 解析中的基础能力之一,主要用于读取独立于页面内容之外的文档级信息。与文本或表格解析不同,元数据解析不依赖页面布局,因此在绝大多数 PDF 文件中都具有较高的稳定性。
在 Java PDF 处理系统中,元数据通常作为前置分析步骤,用于文档分类、流程路由或索引决策。
元数据解析是文档级操作,其实现步骤通常如下:
由于元数据不依赖渲染内容,这一解析过程开销小、速度快,且结果较为一致。
import com.spire.pdf.PdfDocument;
public class parsePdfMetadata {
public static void main(String[] args) {
// 载入 PDF 文档
PdfDocument pdf = new PdfDocument();
pdf.loadFromFile("sample1.pdf");
// 获取 PDF 文档元数据信息
String title = pdf.getDocumentInformation().getTitle();
String author = pdf.getDocumentInformation().getAuthor();
String subject = pdf.getDocumentInformation().getSubject();
String keywords = pdf.getDocumentInformation().getKeywords();
String creator = pdf.getDocumentInformation().getCreator();
String producer = pdf.getDocumentInformation().getProducer();
String creationDate = pdf.getDocumentInformation()
.getCreationDate().toString();
String modificationDate = pdf.getDocumentInformation()
.getModificationDate().toString();
System.out.println(
"Title: " + title +
"\nAuthor: " + author +
"\nSubject: " + subject +
"\nKeywords: " + keywords +
"\nCreator: " + creator +
"\nProducer: " + producer +
"\nCreation Date: " + creationDate +
"\nModification Date: " + modificationDate
);
}
}
控制台输出示例

文档信息字典: 元数据存储在 PDF 的独立结构中,与页面渲染内容无关。
字段完整性: 并非所有 PDF 都包含完整元数据,使用前应进行空值校验。
解析成本低: 不需要遍历页面,适合作为初始解析步骤。
如需访问自定义 PDF 属性,可参考 PdfDocumentInformation API。
由于不受布局和内容流影响,元数据解析在复杂 PDF 中通常比文本或表格解析更加稳定。
在实际项目中,往往需要在同一处理流程中组合多种 PDF 解析能力。
常见的实现模式包括:
将文本、表格、图像和元数据解析视为相互独立但可组合的模块,有助于系统的扩展性、可测试性和长期维护。
即便使用成熟的 Java PDF 解析库,仍存在一些不可避免的限制:
充分理解这些限制,有助于在生产环境中制定合理的解析策略,并降低异常处理复杂度。
在 Java 中进行 PDF 解析时,将其视为一组目标明确、彼此独立的提取操作,往往比线性流程更高效、更可靠。通过分别处理文本、表格和元数据,Java 应用可以稳定地将 PDF 文档转换为可用数据。
借助 Spire.PDF for Java 这样的专业库,开发者可以构建可维护、可扩展,并能满足真实业务需求的 PDF 处理解决方案。
如需全面体验 Spire.PDF for Java 在 PDF 解析方面的能力,可 申请免费试用许可证。
A:可以使用 Spire.PDF for Java 提供的 PdfTextExtractor 与 PdfTextExtractOptions,按页面提取文本,适用于索引、分析或内容迁移场景。
A:通过 PdfTableExtractor 识别表格区域并还原行列结构,解析结果可进一步处理或导出为结构化数据。
A:可以。使用 PdfImageHelper 和 PdfImageInfo 可提取页面中的嵌入图像,也可将整页 PDF 转换为图片。
A:通过 PdfDocumentInformation 获取标题、作者、创建时间等字段,该操作速度快且不依赖页面内容。
A:复杂布局、扫描版 PDF 和自定义字体都会影响解析效果。扫描文档需要先进行 OCR 处理。
Spire.Office for Python 10.12.0 已正式发布。在该版本中,Spire.Doc for Python 优化并增强了 API;Spire.XLS for Python 支持删除 Excel 中的重复行;Spire.Presentation for Pytho 增强了 PPTX 到 PDF 的转换功能;Spire.PDF for Python 支持为数字签名添加时间戳;Spire.Barcode for Python 支持 Linux ARM 平台;Spire.OCR for Python 支持新平台并提升 OCR 识别准确性。此外,一些在转换、处理和保存 Word/Excel/PDF/PowerPoint 文件时出现的问题也已成功被修复。更多新功能及问题修复详情如下。
https://www.e-iceblue.cn/Downloads/Spire-Office-Python.html
调整:
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| Paragraph | GetText | 获取段落的文本内容 |
| Table | SetBorders, ClearBorders |
设置表格边框样式;清除表格的所有边框格式 |
| CellFormat | ClearFormatting | 清除单元格的所有格式 |
| Borders | ClearFormatting, IsShadow |
清除边框格式设置;控制边框是否显示阴影效果 |
| RowFormat | ClearBackground, Height | 清除行背景色;设置行高 |
| StyleCollection | Add (重载) | 增加用于创建样式的重载方法 |
| PreferredWidth | FromPercent, FromPoints |
支持使用百分比或磅值定义宽度 |
| CharacterFormat | LocaleIdBi | 支持双向文本的区域设置 |
| Frame | IsFrame | 判断对象是否为 Frame |
| OfficeMath | ToLaTexMathCode, FromOMMLCode |
将公式对象转换为 LaTeX 数学代码;从 OMML 字符串创建公式对象 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| Chart 及其子对象(包括 ChartAxis、ChartSeries、ChartDataLabelCollection、ChartLegend、ChartTitle 等) | 多个属性和方法 | 支持坐标轴配置、数据标签管理、图例/标题格式设置等 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| CompareOptions | IgnoreTable, IgnoreHeadersAndFooters | 文档比较时忽略表格内容以及页眉/页脚 |
| DifferRevisions | MoveToRevisions, MoveFromRevisions | 获取“移入”和“移出”类型的修订内容 |
| StructureDocumentTag*(包括 Cell / Inline / Row) | RemoveSelfOnly | 仅删除内容控件本身,保留内部内容 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| ToPdfParameterList | PdfImageCompression、DigitalSignatureInfo | 保存到PDF时配置图像压缩以及数字签名信息 |
| MarkdownExportOptions、ListReferences | MoveToRevisions, MoveFromRevisions | 支持 Markdown 导出选项及列表引用 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| ListFormat | ApplyStyle, ApplyListRef | 支持直接应用列表引用及快速样式 |
| ListLevel | Equals, CreatePictureBullet, DeletePictureBullet, PictureBullet | 支持图片项目符号管理及列表级别比较 |
| ListStyle | ListRef, BaseStyle | 支持列表引用及基础样式配置 |
| Document | ListReferences | 获取文档中的列表引用集合 |
新功能:
workbook = Workbook()
workbook.LoadFromFile(inputFile)
sheet = workbook.Worksheets[0]
sheet.RemoveDuplicates()
workbook.SaveToFile(outputFile, FileFormat.Version2013)
workbook.Dispose()
问题修复:
问题修复:
新功能:
doc = PdfDocument ()
doc. LoadFromFile (inputFile)
# Create a digital signature
signature = Security_PdfSignature (doc, doc.Pages.get_Item(0), inputFile_pfx,"08100601", "signature")
signature.SignDetailsFont = PdfFont(PdfFontFamily.Courier,8.0)
# Set the bounds of the signature box
signature.Bounds = RectangleF(PointF(90.0, 550.0), SizeF (180.0, 90.0))
# Configure signature appearance and details
signature.NameLabel = "Digitally signed by:Gary"
signature.LocationInfoLabel ="Location:"
signature.LocationInfo = "CN"
signature.ReasonLabel = "Reaseon:"
signature.Reason = "Ensure authenticity"
signature.ContactInfoLabel = "Contact Number:"
signature.ContactInfo = "028-81705109"
# Set document permissions
signature.DocumentPermissions = PdfCertificationFlags.ForbidChanges.value
# Set graphic mode for the signature
signature.GraphicsMode = Security_GraphicMode.SignImageAndSignDetail
# Set the signature image
signature.SignImageSource = PdfImage.FromFile(inputImage)
#When setting “none", the Image and Detail are distributed on both sides, when setting “Stretch", the image extends to the entire signatu
signature.SignImageLayout = SignImageLayout.none
url = "https://freetsa.org/tsr"
signature.ConfigureTimestamp(url)
signature.ConfigureHttpOCSP (None, None)
signature.Certificated = True
doc.SaveToFile(outputFile)
doc.Close()
pdf = PdfDocument()
pdf.LoadFromFile(inputFile)
textOption = XlsxTextLayoutOptions(True, False, False)
pdf.ConvertOptions.SetPdfToXlsxOptions(textOption)
pdf.SaveToFile(outputFile, FileFormat.XLSX)
pdf.Dispose()
pdf = PdfDocument()
pdf.LoadFromFile(inputFile)
lineOption = XlsxLineLayoutOptions(False,False,False,False)
pdf.ConvertOptions.SetPdfToXlsxOptions(lineOption)
pdf.SaveToFile(outputFile, FileFormat.XLSX)
pdf.Dispose()
# Load the PDF document from the specified input file path
pdf.LoadFromFile(inputFile)
# Set the XlsxSpecialTableLayoutOptions as the conversion options for PDF to XLSX conversion
options = XlsxSpecialTableLayoutOptions(False, False, False)
# Save the PDF document as an Excel file using the specified format and options
pdf.SaveToFile(outputFile, FileFormat.XLSX)
pdf = PdfDocument ()
pdf. LoadFromFile (inputFile)
ofdOptions = OfdOptions()
ofdOptions.UseTempFileStorage = True
pdf.ConvertOptions.SetPdfToOfdOptions(ofdOptions)
pdf.SaveToFile(outputFile,FileFormat.OFD)
# Create an instance of PdfToMarkdownConverter with the input PDF file
converter = PdfToMarkdownConverter(inputFile)
# Configure the converter to skip processing images in the PDF
converter.MarkdownOptions.IgnoreImage = True
# Convert the PDF content to Markdown format and save to the output file
converter.ConvertToMarkdown(outputFile)
converter = PdfToSvgConverter(inputFile)
converter.SvgOptions.ScaleX = 1.0
converter.SvgOptions.ScaleY = 1.0
converter.Convert(outputFile)
问题修复:
新功能:
调整:
优化:
1.支持识别旋转图片。
configureOptions.AutoRotate = True
2.支持按图像中文字的原始位置顺序输出识别结果。
visualText = VisualTextAligner(scanner.Text)
text = visualText.ToString()
我们很高兴地宣布,Spire.Doc 13.12.6 现已发布。本次更新新增了一组文档兼容性相关功能,支持通过指定 Word 版本对文档进行兼容性设置。同时,修复了 Word 转 PDF 过程中页码不正确的问题。更新详情如下:
新功能:
Document doc = new Document();
doc.CompatibilityOptions.UlTrailSpace = false;
doc.CompatibilityOptions.AdjustLineHeightInTable = true;
doc.CompatibilityOptions.SpaceForUL = true;
doc.CompatibilityOptions.ApplyBreakingRules = true;
doc.CompatibilityOptions.DoNotExpandShiftReturn = false;
doc.CompatibilityOptions.OverrideTableStyleFontSizeAndJustification = false;
doc.CompatibilityOptions.DoNotAutofitConstrainedTables = true;
doc.SaveToFile("outputFile");
Document doc = new Document();
doc.LoadFromFile("inputtFile");
Spire.Doc.Settings.CompatibilityOptions options = doc.CompatibilityOptions;
Document doc = new Document();
doc.LoadFromFile(inputFile);
// Set properties
doc.CompatibilityOptions.UlTrailSpace = false;
doc.CompatibilityOptions.AdjustLineHeightInTable = true;
doc.CompatibilityOptions.SpaceForUL = true;
doc.CompatibilityOptions.ApplyBreakingRules = true;
doc.CompatibilityOptions.DoNotExpandShiftReturn = false;
doc.CompatibilityOptions.OverrideTableStyleFontSizeAndJustification = false;
doc.CompatibilityOptions.DoNotAutofitConstrainedTables = true;
// Set FileFormat when saving to preserve effects
doc.SaveToFile(outputFile_after, FileFormat.Docx2016);
// Using version compatibility will reset previously set properties
Spire.Doc.Settings.CompatibilityOptions options = doc.CompatibilityOptions;
doc.CompatibilityOptions.OptimizeForWordVersion(WordVersion.Word2016);
PrintCompatibilityOptions(options, outputFile);
doc.Close();
问题修复:
Spire.PDF for Python 11.12.1 现已正式发布,本次更新带来了多项新功能,包括为数字签名添加时间戳、配置 PDF 转 Excel 时的多种布局选项、以及在 PDF 转 Markdown 时忽略图像。同时,本版本还修复了两个已知问题。更多详情如下。
新功能:
doc = PdfDocument ()
doc. LoadFromFile (inputFile)
# Create a digital signature
signature = Security_PdfSignature (doc, doc.Pages.get_Item(0), inputFile_pfx,"08100601", "signature")
signature.SignDetailsFont = PdfFont(PdfFontFamily.Courier,8.0)
# Set the bounds of the signature box
signature.Bounds = RectangleF(PointF(90.0, 550.0), SizeF (180.0, 90.0))
# Configure signature appearance and details
signature.NameLabel = "Digitally signed by:Gary"
signature.LocationInfoLabel ="Location:"
signature.LocationInfo = "CN"
signature.ReasonLabel = "Reaseon:"
signature.Reason = "Ensure authenticity"
signature.ContactInfoLabel = "Contact Number:"
signature.ContactInfo = "028-81705109"
# Set document permissions
signature.DocumentPermissions = PdfCertificationFlags.ForbidChanges.value
# Set graphic mode for the signature
signature.GraphicsMode = Security_GraphicMode.SignImageAndSignDetail
# Set the signature image
signature.SignImageSource = PdfImage.FromFile(inputImage)
#When setting “none", the Image and Detail are distributed on both sides, when setting “Stretch", the image extends to the entire signatu
signature.SignImageLayout = SignImageLayout.none
url = "https://freetsa.org/tsr"
signature.ConfigureTimestamp(url)
signature.ConfigureHttpOCSP (None, None)
signature.Certificated = True
doc.SaveToFile(outputFile)
doc.Close()
pdf = PdfDocument()
pdf.LoadFromFile(inputFile)
textOption = XlsxTextLayoutOptions(True, False, False)
pdf.ConvertOptions.SetPdfToXlsxOptions(textOption)
pdf.SaveToFile(outputFile, FileFormat.XLSX)
pdf.Dispose()
pdf = PdfDocument()
pdf.LoadFromFile(inputFile)
lineOption = XlsxLineLayoutOptions(False,False,False,False)
pdf.ConvertOptions.SetPdfToXlsxOptions(lineOption)
pdf.SaveToFile(outputFile, FileFormat.XLSX)
pdf.Dispose()
# Load the PDF document from the specified input file path
pdf.LoadFromFile(inputFile)
# Set the XlsxSpecialTableLayoutOptions as the conversion options for PDF to XLSX conversion
options = XlsxSpecialTableLayoutOptions(False, False, False)
# Save the PDF document as an Excel file using the specified format and options
pdf.SaveToFile(outputFile, FileFormat.XLSX)
pdf = PdfDocument ()
pdf. LoadFromFile (inputFile)
ofdOptions = OfdOptions()
ofdOptions.UseTempFileStorage = True
pdf.ConvertOptions.SetPdfToOfdOptions(ofdOptions)
pdf.SaveToFile(outputFile,FileFormat.OFD)
# Create an instance of PdfToMarkdownConverter with the input PDF file
converter = PdfToMarkdownConverter(inputFile)
# Configure the converter to skip processing images in the PDF
converter.MarkdownOptions.IgnoreImage = True
# Convert the PDF content to Markdown format and save to the output file
converter.ConvertToMarkdown(outputFile)
converter = PdfToSvgConverter(inputFile)
converter.SvgOptions.ScaleX = 1.0
converter.SvgOptions.ScaleY = 1.0
converter.Convert(outputFile)
问题修复:
Spire.OCR for Python 1.9.13 已发布。本次版本先进行了若干依赖与平台方面的调整,随后增强了错误处理与 OCR 识别能力。详细更新如下:
调整:
优化:
1.支持识别旋转图片。
configureOptions.AutoRotate = True
2.支持按图像中文字的原始位置顺序输出识别结果。
visualText = VisualTextAligner(scanner.Text)
text = visualText.ToString()
随着企业数字化转型加速,纯前端文档处理已成为 Web 应用的核心需求之一。Angular 作为成熟的企业级前端框架,以其强类型校验、组件化架构和高效的状态管理能力,广泛应用于 OA 系统、文档平台、教育管理系统等复杂场景。
本文将详细介绍如何将 Spire.OfficeJS 集成到 Angular 项目中,实现本地文件上传、在线编辑、格式转换与下载等核心功能。
内容概览
Spire.OfficeJS 是一款企业级在线文档处理与编辑解决方案,包含 Spire.WordJS、Spire.ExcelJS、Spire.PresentationJS、Spire.PDFJS 四个模块。该产品无需依赖本地 Office 软件,也无需安装任何插件,即可在浏览器中实现 Word、Excel、PPT、PDF 等主流格式文件的在线预览、实时编辑、批注、与格式转换等功能,同时具备云原生、跨平台、高安全性等企业级特性。
核心优势包括:
在开始集成前,需先完成以下准备工作,确保开发环境符合要求。
node -v
npm -v

Angular CLI 是快速构建 Angular 项目的工具,通过 npm 全局安装:
npm install -g @angular/cli
ng version

通过 Angular CLI 创建基础项目,为后续集成 Spire.OfficeJS 搭建框架。
F:\angular):ng new spireOfficeJS --skip-git


使用 VS Code 打开项目,在终端执行以下命令启动开发服务器:
npm run start
浏览器打开 http://localhost:4200/,若显示 “Hello, spireOfficeJS” 及 Angular 默认页面,则项目初始化成功。
下载 Spire.OfficeJS 产品包,解压并运行 run_genallfonts 脚本后,需将更新后的产品包中的web 文件夹复制到 Angular 项目的静态目录中,以确保编辑器脚本可被正常访问。具体操作如下:

public/spire.cloud)。public/spire.cloud/ 目录下。public/spire.cloud/web/editors/spireapi/SpireCloudEditor.js。
⚠️ 注意:此路径需与后续配置的 office-js.ts 中路径完全一致,否则编辑器无法加载
用于同步文件上传数据(文件对象 + 二进制数据),确保编辑器组件可访问。
(1)安装依赖
使用 NgRx Signals 管理文件数据,在VS Code 终端执行以下安装命令:
npm install @ngrx/effects @ngrx/signals @ngrx/store
(2)创建状态管理 Store
import { signalStore, withState, withMethods, patchState } from '@ngrx/signals';
// 定义文件状态接口
interface FileState {
file: File | null; // 上传的文件对象
fileUint8Data: Uint8Array | null; // 文件二进制数据
}
// 初始状态
const initialState: FileState = {
file: null,
fileUint8Data: null,
};
// 创建全局 Store
export const fileStore = signalStore(
{ providedIn: 'root' },
withState(initialState),
withMethods((store) => ({
// 更新文件对象
setFileData(data: File | null): void {
patchState(store, { file: data });
},
// 更新文件二进制数据
setFileUint8Data(uint8Data: Uint8Array | null): void {
patchState(store, { fileUint8Data: uint8Data });
}
}))
);
(1)创建组件
在终端执行以下命令,创建文件上传组件和编辑器集成组件:
ng g c spire/uploadFile
ng g c spire/officeJS
执行完成后,src/app/spire/ 目录下会生成 upload-file/ 和 office-js/ 两个组件文件夹。

(2)文件上传组件(upload-file)配置
上传组件负责接收用户拖放或选择的文件,转换为二进制格式后存储到全局状态,再跳转至编辑器页面。
upload-file.css)::host {
display: block;
min-height: 100vh;
}
.upload-main {
min-height: 100vh;
display: flex;
justify-content: center;
align-items: center;
background-color: #f5f5f5;
}
.upload-container {
width: 80%;
max-width: 600px;
padding: 40px;
background-color: white;
border-radius: 8px;
box-shadow: 0 4px 12px rgba(0, 0, 0, 0.1);
text-align: center;
}
.drop-area {
border: 2px dashed #ccc;
border-radius: 6px;
padding: 40px;
margin-bottom: 20px;
transition: all 0.3s;
}
.drop-area.highlight {
border-color: #4CAF50;
background-color: #f0fff0;
}
button {
background-color: #4CAF50;
color: white;
border: none;
padding: 10px 20px;
border-radius: 4px;
cursor: pointer;
font-size: 16px;
margin-top: 10px;
}
button:hover {
background-color: #45a049;
}
#fileInput {
display: none; /* 隐藏原生文件选择框 */
}
upload-file.html),支持拖放上传和点击选择文件两种方式:<main class="upload-main">
<div class="upload-container">
<h2>拖放文件上传</h2>
<div class="drop-area" id="dropArea">
<p>拖放文件到浏览器</p>
<p>或</p>
<button id="browseBtn" #browseBtn (click)="handleButtonClick($event)">选择文件</button>
<input type="file" id="fileInput" #fileInput (change)="handleDrop($event)">
</div>
</div>
</main>
upload-file.ts),处理文件拖放、选择、二进制转换与状态储存:import { ViewChild, ElementRef, Component, AfterViewInit, inject } from '@angular/core';
import { Router } from '@angular/router';
import { fileStore } from '../../store/index';
@Component({
selector: 'app-upload-file',
imports: [],
templateUrl: './upload-file.html',
styleUrl: './upload-file.css',
})
export class UploadFile implements AfterViewInit {
constructor(private router: Router) { }
// 注入状态管理 Store
store = inject(fileStore);
// 绑定 HTML 元素
@ViewChild('browseBtn') browseBtn!: ElementRef<HTMLButtonElement>;
@ViewChild('fileInput') fileInput!: ElementRef<HTMLInputElement>;
file: File | null = null; // 上传的文件对象
fileUint8Data: Uint8Array | null = null; // 文件二进制数据
// 组件视图初始化完成后执行
ngAfterViewInit() {
this.init();
}
// 初始化拖放事件监听
init() {
// 阻止浏览器默认拖放行为
['dragenter', 'dragover', 'dragleave', 'drop'].forEach(eventName => {
document.addEventListener(eventName, this.preventDefaults, false);
});
// 监听文件拖放事件
document.addEventListener('drop', (e) => {
this.handleDrop.call(this, e);
}, false);
}
// 阻止默认事件
preventDefaults(e: Event) {
e.preventDefault();
e.stopPropagation();
}
// 点击“选择文件”按钮触发原生文件选择框
handleButtonClick(e: Event) {
e.preventDefault();
this.fileInput.nativeElement.click();
}
// 处理文件拖放/选择事件
async handleDrop(e: any) {
// 获取文件对象(拖放或点击选择)
if (e.target && e.target.files) {
this.file = e.target.files[0];
} else if (e.dataTransfer && e.dataTransfer.files) {
this.file = e.dataTransfer.files[0];
}
// 转换文件为 Uint8Array 二进制格式
this.fileUint8Data = await this.handleFile(this.file) as Uint8Array;
// 更新状态到 Store(供编辑器组件使用)
this.store.setFileData(this.file);
this.store.setFileUint8Data(this.fileUint8Data);
// 跳转到编辑器页面
this.openDocument();
}
// 将文件转换为 Uint8Array 二进制数据
handleFile(file: any) {
return new Promise((resolve, reject) => {
const reader = new FileReader();
reader.onload = () => {
const arrayBuffer = reader.result as ArrayBuffer;
const uint8Array = new Uint8Array(arrayBuffer);
resolve(uint8Array);
};
reader.onerror = (error) => reject(error);
reader.readAsArrayBuffer(file); // 以 ArrayBuffer 格式读取文件
});
}
// 跳转到编辑器页面(路由导航)
openDocument() {
this.router.navigate(['spire']);
}
}
(3) 编辑器组件(office-js)配置
编辑器组件负责加载 Spire.OfficeJS 脚本、初始化编辑器实例、配置编辑权限与功能,是核心交互模块。
office-js.html):<div class="form">
<div id="iframeEditor">
</div>
</div>
office-js.ts):import { Component, AfterViewInit, inject } from '@angular/core';
import { Router } from '@angular/router';
import { fileStore } from '../../store/index';
// 声明 SpireCloudEditor 全局变量(来自产品脚本)
declare const SpireCloudEditor: any;
@Component({
selector: 'app-office-js',
imports: [],
templateUrl: './office-js.html',
styleUrl: './office-js.css',
})
export class OfficeJS implements AfterViewInit {
constructor(private router: Router) { };
// 注入状态管理 Store
store = inject(fileStore);
// 从 Store 获取文件数据
file = this.store.file() as File;
fileUint8Data = this.store.fileUint8Data() as Uint8Array;
originUrl = window.location.origin; // 当前项目域名
Editor: any; // 编辑器实例
config: any; // 编辑器配置
Api: any; // 编辑器 API
// 组件视图初始化完成后执行
ngAfterViewInit() {
this.init();
}
// 初始化校验(无文件则跳回上传页)
init() {
if (!this.file) {
this.router.navigate(['']); // 跳转到上传页面
return;
}
this.loadSrcipt(); // 加载编辑器脚本
}
// 动态加载 SpireCloudEditor.js 脚本
loadSrcipt() {
const script = document.createElement('script');
// 脚本路径需与静态资源部署路径一致
script.setAttribute('src', '/spire.cloud/web/editors/spireapi/SpireCloudEditor.js');
script.onload = () => this.initEditor(); // 脚本加载完成后初始化编辑器
document.head.appendChild(script);
}
// 初始化编辑器配置与实例
initEditor() {
const iframeId = 'iframeEditor'; // 与模板文件中容器 ID 一致
this.initConfig(); // 配置编辑器参数
// 创建编辑器实例
this.Editor = new SpireCloudEditor.OpenApi(iframeId, this.config);
this.Api = this.Editor.GetOpenApi(); // 获取编辑器 API (用于扩展功能)
this.OnWindowReSize(); // 适配窗口大小
}
// 编辑器核心配置(文件信息 + 用户权限 + 编辑器行为)
initConfig() {
this.config = {
"fileAttrs": {
"fileInfo": {
"name": this.file.name, // 文件名
"ext": this.getFileExtension(), // 文件后缀
"primary": String(new Date().getTime()), // 唯一标识(时间戳)
"creator": "",
"createTime": ""
},
"sourceUrl": `${this.originUrl}/files/__ffff_192.168.3.121/${this.file.name}`,
"createUrl": `${this.originUrl}/open`,
"mergeFolderUrl": "",
"fileChoiceUrl": "",
"templates": {}
},
"user": {
"id": "uid-1",
"name": "Jonn",
"canSave": true, // 允许保存文件
},
"editorAttrs": {
"editorMode": this.file.name.endsWith('.pdf') ? 'view' : "edit", // PDF 默认为预览模式
"editorWidth": "100%", // 编辑器宽度
"editorHeight": "100%", // 编辑器高度
"editorType": "document", // 编辑器类型(文档)
"platform": "desktop", // 平台类型(桌面端)
"viewLanguage": "zh", // 界面语言(中文)
"isReadOnly": false, // 非只读模式
"canChat": true, // 启用聊天功能
"canComment": true, // 启用批注功能
"canReview": true, // 启用审阅功能
"canDownload": true, // 允许下载文件
"canEdit": this.file.name.endsWith('.pdf') ? false : true, // PDF 禁止编辑
"canForcesave": true, // 允许强制保存
"embedded": {
"saveUrl": "",
"embedUrl": "",
"shareUrl": "",
"toolbarDocked": "top" // 工具栏置顶
},
// 启用 WebAssembly 加速(提升编辑和转换性能)
"useWebAssemblyDoc": true,
"useWebAssemblyExcel": true,
"useWebAssemblyPpt": true,
"useWebAssemblyPdf": true,
// 许可证配置(若有许可证,填写对应密钥)
"spireDocJsLicense": "",
"spireXlsJsLicense": "",
"spirePresentationJsLicense": "",
"spirePdfJsLicense": "",
"serverless": {
"useServerless": true,
"baseUrl": this.originUrl,
"fileData": this.fileUint8Data, // 核心:传入文件二进制数据
},
"events": {
"onSave": this.onFileSave // 保存回调事件
},
"plugins": {
"pluginsData": []
}
}
};
}
// 窗口大小适配
OnWindowReSize() {
const wrapEl = document.getElementsByClassName("form") as any;
if (wrapEl.length) {
wrapEl[0].style.height = window.innerHeight + "px";
window.scrollTo(0, -1);
}
}
// 获取文件后缀名
getFileExtension() {
const filename = this.file.name.split(/[\\/]/).pop() as String;
return filename.substring(filename.lastIndexOf('.') + 1).toLowerCase() || '';
}
// 自定义保存逻辑(可根据需求扩展)
onFileSave(data: any) {
console.log('保存的数据:', data);
}
}
为上传页面和编辑器页面配置路由,实现页面跳转。
app.routes.ts),配置上传页和编辑器页的路由:import { Routes } from '@angular/router';
import { UploadFile } from './spire/upload-file/upload-file';
import { OfficeJS } from './spire/office-js/office-js';
export const routes: Routes = [
{ path: '', component: UploadFile }, // 默认路由:文件上传页
{ path: 'spire', component: OfficeJS } // 编辑器路由:/spire
];
app.config.ts),确保路由配置生效:import { ApplicationConfig, provideBrowserGlobalErrorListeners, provideZoneChangeDetection } from '@angular/core';
import { provideRouter } from '@angular/router';
import { routes } from './app.routes';
export const appConfig: ApplicationConfig = {
providers: [
provideBrowserGlobalErrorListeners(),
provideZoneChangeDetection({ eventCoalescing: true }),
provideRouter(routes), // 注入路由
]
};
app.html),确保页面正确渲染:<main class="app-main">
<router-outlet /> <!-- 路由出口:渲染当前路由对应的组件 -->
</main>
保存所有修改后,在项目根目录(F:\angular\spireOfficeJS)执行以下命令,重启开发服务器:
npm run start



Q1:编辑器加载失败,页面空白?
SpireCloudEditor.js 路径与 office-js.ts 中 script.src 一致Q2:安装 NgRx 依赖时提示 peer dependency 冲突?
npm install @ngrx/signals @ngrx/store --legacy-peer-deps
Q3:启动项目时,浏览器控制台报错,提示找不到 zone.js 模块?
zone.js 依赖),但 Spire.OfficeJS 依赖 zone.js 处理异步事件。npm install zone.js --savesrc/main.ts,在文件顶部添加导入:import 'zone.js';下载 Angular 集成 Spire.OfficeJS 的完整示例项目,包含所有配置文件与代码,直接运行即可体验功能。
如需去除水印或解锁全部功能,可联系我们获取 30 天临时许可证。
Spire.Doc 13.12.2 现已正式发布。本版本支持双行合一功能,强化了 Word 转 PDF 的效果。同时,支持设置段落文本的“Horizontal in Vertical”属性、Markdown 转 Docx 时从模板文档复制样式,以及获取样式更改修订。更多详情如下。
新功能:
Document doc = new Document();
Section section = doc.AddSection();
Spire.Doc.Documents.Paragraph paragraph = section.AddParagraph();
Spire.Doc.Fields.TextRange farEastLayout = paragraph.AppendText("test");
FarEastLayout style = new FarEastLayout();
style.Vertical = true;
farEastLayout.CharacterFormat.FarEastLayout = style;
doc.SaveToFile(outputFile, FileFormat.Docx);
doc.Close();
//Load template documents with existing styles
Document temple = new Document();
temple.LoadFromFile("temple.docx");
//Load markdown file
Document doc = new Document();
doc = new Document(@"Doc.md");
//Copy styles from template documents
doc.CopyStylesFromTemplate(temple);
//Save
doc.SaveToFile(@"Doc.docx", Spire.Doc.FileFormat.Docx2016);
Document doc = new Document();
doc.LoadFromFile(inputFile);
RevisionInfoCollection revisionInfoCollection = doc.GetRevisionInfos();
StringBuilder sb = new StringBuilder();
foreach (RevisionInfo revisionInfo in revisionInfoCollection)
{
if (revisionInfo.RevisionType == RevisionType.FormatChange)
{
if (revisionInfo.OwnerObject is Spire.Doc.Fields.TextRange)
{
TextRange range = (TextRange)revisionInfo.OwnerObject;
sb.AppendLine("TextRange:" + range.Text + "\r\n");
doc.RevisionsView = RevisionsView.Original;
sb.AppendLine("Original style:" + "isBold:" + range.CharacterFormat.Bold + ";" + "TextColor:" + range.CharacterFormat.TextColor + ";HighlightColor:" + range.CharacterFormat.HighlightColor + ";FontName:" + range.CharacterFormat.FontName + ";UnderlineStyle:" + range.CharacterFormat.UnderlineStyle + "\r\n");
doc.RevisionsView = RevisionsView.Final;
sb.AppendLine("Final style:" + "isBold:" + range.CharacterFormat.Bold + ";" + "TextColor:" + range.CharacterFormat.TextColor + ";HighlightColor:" + range.CharacterFormat.HighlightColor + ";FontName:" + range.CharacterFormat.FontName + ";UnderlineStyle:" + range.CharacterFormat.UnderlineStyle + "\r\n");
}
}
}
File.WriteAllText(outputFile, sb.ToString());
doc.Close();
问题修复:
Spire.Doc for Java 13.12.2 现已正式发布。该版本支持检测写保护密码是否正确,此外还修复了一些在转换 Word 到 PDF、HTML 到 Word,和替换书签时出现的问题。详情如下。
新功能:
Boolean protectionPassword = document.checkWriteProtectionPassword("password");
问题修复:
Spire.Doc for Python 13.12.0 已正式发布,该版本带来了多项重要的 API 增强功能,包括针对文档元素的精细化格式控制、图表配置选项的强化以及文档比较功能的升级。此外,还对辅助功能进行了增强,重构了列表系统,并对整体 API 进行了全面优化,从而显著提升了用户体验。具体更新内容如下。
调整:
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| Paragraph | GetText | 获取段落的文本内容 |
| Table | SetBorders, ClearBorders |
设置表格边框样式;清除表格的所有边框格式 |
| CellFormat | ClearFormatting | 清除单元格的所有格式 |
| Borders | ClearFormatting, IsShadow |
清除边框格式设置;控制边框是否显示阴影效果 |
| RowFormat | ClearBackground, Height | 清除行背景色;设置行高 |
| StyleCollection | Add (重载) | 增加用于创建样式的重载方法 |
| PreferredWidth | FromPercent, FromPoints |
支持使用百分比或磅值定义宽度 |
| CharacterFormat | LocaleIdBi | 支持双向文本的区域设置 |
| Frame | IsFrame | 判断对象是否为 Frame |
| OfficeMath | ToLaTexMathCode, FromOMMLCode |
将公式对象转换为 LaTeX 数学代码;从 OMML 字符串创建公式对象 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| Chart 及其子对象(包括 ChartAxis、ChartSeries、ChartDataLabelCollection、ChartLegend、ChartTitle 等) | 多个属性和方法 | 支持坐标轴配置、数据标签管理、图例/标题格式设置等 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| CompareOptions | IgnoreTable, IgnoreHeadersAndFooters | 文档比较时忽略表格内容以及页眉/页脚 |
| DifferRevisions | MoveToRevisions, MoveFromRevisions | 获取“移入”和“移出”类型的修订内容 |
| StructureDocumentTag*(包括 Cell / Inline / Row) | RemoveSelfOnly | 仅删除内容控件本身,保留内部内容 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| ToPdfParameterList | PdfImageCompression、DigitalSignatureInfo | 保存到PDF时配置图像压缩以及数字签名信息 |
| MarkdownExportOptions、ListReferences | MoveToRevisions, MoveFromRevisions | 支持 Markdown 导出选项及列表引用 |
| 类名 | 新功能 | 功能说明 |
|---|---|---|
| ListFormat | ApplyStyle, ApplyListRef | 支持直接应用列表引用及快速样式 |
| ListLevel | Equals, CreatePictureBullet, DeletePictureBullet, PictureBullet | 支持图片项目符号管理及列表级别比较 |
| ListStyle | ListRef, BaseStyle | 支持列表引用及基础样式配置 |
| Document | ListReferences | 获取文档中的列表引用集合 |