Spire.Doc 12.2.10已发布。该版本支持解析Word文档中的GIF格式内容。此外,还修复了一些已知问题,如修复了获取出的项目符号不正确的问题。详情请阅读以下内容。
新功能:
问题修复:
在 Word 文档中进行页面的添加、插入和删除是管理和展示内容的关键步骤。通过添加或插入新页面,您可以扩展文档以容纳更多内容,使其更有条理和易读性。删除页面则有助于简化文档,去除不必要或错误的信息。这些操作可以提升文档的整体质量和清晰度。本文将介绍如何使用 Spire.Doc for Java 在 Java 项目中添加、插入和删除 Word 文档中的页面。
首先,您需要在 Java 程序中添加 Spire.Doc.jar 文件作为依赖项。您可以从此链接下载 JAR 文件;如果您使用 Maven,则可以通过在 pom.xml 文件中添加以下代码导入 JAR 文件。
<repositories>
<repository>
<id>com.e-iceblue</id>
<name>e-iceblue</name>
<url>https://repo.e-iceblue.cn/repository/maven-public/</url>
</repository>
</repositories>
<dependencies>
<dependency>
<groupId>e-iceblue</groupId>
<artifactId>spire.doc</artifactId>
<version>14.9.5</version>
</dependency>
</dependencies>
在 Word 文档末尾添加新页面的步骤包括定位到最后一个章节,然后在该章节的末尾段落插入分页符。这种方法可以确保后续添加的内容会从新页面开始显示,保持文档结构的清晰和连贯。以下是详细步骤:
import com.spire.doc.*;
import com.spire.doc.documents.*;
public class AddOnePage {
public static void main(String[] args) {
// 创建一个新的文档对象
Document document = new Document();
// 从文件加载示例文档
document.loadFromFile("示例.docx");
// 获取文档的最后一个章节的正文部分
Body body = document.getLastSection().getBody();
// 在正文最后一个段落后插入分页符
body.getLastParagraph().appendBreak(BreakType.Page_Break);
// 创建一个新的段落样式
ParagraphStyle paragraphStyle = new ParagraphStyle(document);
paragraphStyle.setName("CustomParagraphStyle1");
paragraphStyle.getParagraphFormat().setLineSpacing(12);
paragraphStyle.getParagraphFormat().setAfterSpacing(8);
paragraphStyle.getCharacterFormat().setFontName("微软雅黑");
paragraphStyle.getCharacterFormat().setFontSize(12);
// 将段落样式添加到文档的样式集合中
document.getStyles().add(paragraphStyle);
// 创建新的段落并设置文本内容
Paragraph paragraph = new Paragraph(document);
paragraph.appendText("非常感谢您使用我们的 Spire.Doc for Java 产品。试用版除了会在生成的结果文档中添加红色水印,而且仅支持转换前 10 页到其它格式。当您购买并应用 license 后,会成功移除这些水印信息并解除功能限制。");
// 应用段落样式
paragraph.applyStyle(paragraphStyle.getName());
// 将段落添加到正文的内容集合中
body.getChildObjects().add(paragraph);
// 创建另一个新的段落并设置文本内容
paragraph = new Paragraph(document);
paragraph.appendText("为了更完整的试用我们的产品,我们免费提供一个月临时 license 给我们的每一位客户。请发送邮件到 sales @e-iceblue.com,我们会在一个工作日内将license发送给您。");
// 应用段落样式
paragraph.applyStyle(paragraphStyle.getName());
// 将段落添加到正文的内容集合中
body.getChildObjects().add(paragraph);
// 保存文档到指定路径
document.saveToFile("添加一个页面.docx", FileFormat.Docx);
// 关闭文档
document.close();
// 释放文档对象的资源
document.dispose();
}
}
在插入新页面之前,需要确定指定页面内容在节中的结束位置索引,然后逐个将新页面的内容添加到文档中。为了确保内容与后续页面分隔开来,需要在适当位置插入分页符。详细步骤如下:
import com.spire.doc.*;
import com.spire.doc.pages.*;
import com.spire.doc.documents.*;
public class InsertOnePage {
public static void main(String[] args) {
// 创建一个新的文档对象
Document document = new Document();
// 从文件加载示例文档
document.loadFromFile("示例.docx");
// 创建固定布局文档对象
FixedLayoutDocument layoutDoc = new FixedLayoutDocument(document);
// 获取第一页
FixedLayoutPage page = layoutDoc.getPages().get(0);
// 获取文档正文部分
Body body = page.getSection().getBody();
// 获取当前页最后一列的段落
Paragraph paragraphEnd = page.getColumns().get(0).getLines().getLast().getParagraph();
// 初始化结束索引
int endIndex = 0;
if (paragraphEnd != null)
{
// 获取最后一个段落的索引
endIndex = body.getChildObjects().indexOf(paragraphEnd);
}
// 创建一个新的段落样式
ParagraphStyle paragraphStyle = new ParagraphStyle(document);
paragraphStyle.setName("CustomParagraphStyle1");
paragraphStyle.getParagraphFormat().setLineSpacing(12);
paragraphStyle.getParagraphFormat().setAfterSpacing(8);
paragraphStyle.getCharacterFormat().setFontName("微软雅黑");
paragraphStyle.getCharacterFormat().setFontSize(12);
// 将样式添加到文档中
document.getStyles().add(paragraphStyle);
// 创建新的段落并设置文本内容
Paragraph paragraph = new Paragraph(document);
paragraph.appendText("非常感谢您使用我们的 Spire.Doc for Java 产品。试用版除了会在生成的结果文档中添加红色水印,而且仅支持转换前 10 页到其它格式。当您购买并应用 license 后,会成功移除这些水印信息并解除功能限制。");
// 应用段落样式
paragraph.applyStyle(paragraphStyle.getName());
// 在指定位置插入段落
body.getChildObjects().insert(endIndex + 1, paragraph);;
// 创建另一个新的段落并设置文本内容
paragraph = new Paragraph(document);
paragraph.appendText("为了更完整的试用我们的产品,我们免费提供一个月临时 license 给我们的每一位客户。请发送邮件到 sales @e-iceblue.com,我们会在一个工作日内将license发送给您。");
// 应用段落样式
paragraph.applyStyle(paragraphStyle.getName());
// 添加分页符
paragraph.appendBreak(BreakType.Page_Break);
// 在指定位置插入段落
body.getChildObjects().insert(endIndex + 2, paragraph);
// 保存文档到指定路径
document.saveToFile("在指定的页面后插入新的一页.docx",FileFormat.Docx);
// 关闭并释放文档对象的资源
document.close();
document.dispose();
}
}
要删除一个页面的内容,首先需要找到该页面的起始元素和结束元素在文档中的位置索引。接着,通过循环逐个移除这些元素,实现删除整个页面的内容。详细步骤如下:
import com.spire.doc.*;
import com.spire.doc.pages.*;
import com.spire.doc.documents.*;
public class RemoveOnePage {
public static void main(String[] args) {
// 创建一个新的文档对象
Document document = new Document();
// 从文件加载示例文档
document.loadFromFile("示例.docx");
// 创建固定布局文档对象
FixedLayoutDocument layoutDoc = new FixedLayoutDocument(document);
// 获取第二页
FixedLayoutPage page = layoutDoc.getPages().get(1);;
// 获取页面的节
Section section = page.getSection();
// 获取第一页第一列的段落
Paragraph paragraphStart = page.getColumns().get(0).getLines().getFirst().getParagraph();
int startIndex = 0;
if (paragraphStart != null)
{
// 获取起始段落的索引
startIndex = section.getBody().getChildObjects().indexOf(paragraphStart);
}
// 获取最后一页最后一列的段落
Paragraph paragraphEnd = page.getColumns().get(0).getLines().getLast().getParagraph();
int endIndex = 0;
if (paragraphEnd != null)
{
// 获取结束段落的索引
endIndex = section.getBody().getChildObjects().indexOf(paragraphEnd);
}
// 删除指定范围内的段落
for (int i = 0; i <= (endIndex - startIndex); i++)
{
section.getBody().getChildObjects().removeAt(startIndex);
}
// 保存文档到指定路径
document.saveToFile("删除一个页面.docx", FileFormat.Docx);
// 关闭并释放文档对象的资源
document.close();
document.dispose();
}
}
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
添加 OLE 对象到 PowerPoint 中的好处主要包括集成多种数据类型、增强交互性和动态性、提高工作效率、保持内容更新、增强视觉效果以及简化复杂流程。这些好处使得 OLE 对象成为 PowerPoint 中一种强大且灵活的工具,有助于创建更加生动、专业且高效的演示文稿。在本文中,我们将详细介绍如何使用 Spire.Presentation for Python 在 Python 中向 PowerPoint 演示文稿插入、提取或修改 OLE 对象。
本教程需要用到 Spire.Presentation for Python 和 plum-dispatch v1.7.4。可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.Presentation如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.Presentation for Python
Spire.Presentation for Python 提供了 IShape.AppendOleObject() 方法,用于在 PowerPoint 演示文稿中插入 OLE 对象。详细步骤如下:
from spire.presentation.common import *
from spire.presentation import *
# 创建一个新的 PowerPoint 演示文稿对象
ppt = Presentation()
slide = ppt.Slides[0]
# 从流中添加 Excel 图像
excelImageStream = Stream("excel.png")
oleImage = ppt.Images.AppendStream(excelImageStream)
excelImageStream.Close()
# 设置位置并将Excel文件添加到OLE中
excelRec = RectangleF.FromLTRB(100, 60, oleImage.Width+100, oleImage.Height+60)
oleStream = Stream("Excel文件.xlsx")
oleObject = slide.Shapes.AppendOleObject("excel", oleStream, excelRec)
oleObject.SubstituteImagePictureFillFormat.Picture.EmbedImage = oleImage
oleObject.ProgId = "Excel.Sheet.12"
# 从流中添加 Zip 图像
zipImageStream = Stream("zip.png")
zipOleImage = ppt.Images.AppendStream(zipImageStream)
zipImageStream.Close()
# 设置位置并将Zip文件添加到OLE中
zipRec = RectangleF.FromLTRB(100, oleImage.Height+100, zipOleImage.Width+100, zipOleImage.Height+oleImage.Height+100)
zipOleStream = Stream("展示PPT.zip")
zipOleObject = slide.Shapes.AppendOleObject("zipPackage", zipOleStream, zipRec)
zipOleObject.ProgId = "Package"
zipOleObject.SubstituteImagePictureFillFormat.Picture.EmbedImage = zipOleImage
# 关闭流
oleStream.Close()
zipOleStream.Close()
# 保存 PowerPoint 演示文稿
ppt.SaveToFile("添加OLE对象.pptx", FileFormat.Pptx2010)
ppt.Dispose()
如果您喜欢 PowerPoint 演示文稿中嵌入的 OLE 对象,并想在其他地方使用它们,可以将它们提取出来并保存到指定的磁盘中。以下步骤详细演示了如何从 PowerPoint 中提取OLE对象:
from spire.presentation.common import *
from spire.presentation import *
# 设置输出文件路径
outputFile_px = "提取OLE/ExtractOLEObject.pptx"
outputFile_p = "提取OLE/ExtractOLEObject.ppt"
outputFile_xls = "提取OLE/ExtractOLEObject.xls"
outputFile_xlsx = "提取OLE/ExtractOLEObject.xlsx"
outputFile_doc = "提取OLE/ExtractOLEObject.doc"
outputFile_docx = "提取OLE/ExtractOLEObject.docx"
outputFile_zip = "提取OLE/ExtractOLEObject.zip"
# 创建 PowerPoint 演示文稿对象
presentation = Presentation()
# 加载带有 OLE 对象的 PowerPoint 演示文稿
presentation.LoadFromFile("提取Ole.pptx")
# 遍历每一页的Shape
for slide in presentation.Slides:
for shape in slide.Shapes:
# 检查Shape是否为 OLE 对象
if isinstance(shape, IOleObject):
oleObject = shape if isinstance(shape, IOleObject) else None
stream = oleObject.Data
# 根据不同的 ProgId 保存不同类型的文件
if oleObject.ProgId == "Excel.Sheet.8":
stream.Save(outputFile_xls)
elif oleObject.ProgId == "Excel.Sheet.12":
stream.Save(outputFile_xlsx)
elif oleObject.ProgId == "Word.Document.8":
stream.Save(outputFile_doc)
elif oleObject.ProgId == "Word.Document.12":
stream.Save(outputFile_docx)
elif oleObject.ProgId == "PowerPoint.Show.8":
stream.Save(outputFile_p)
elif oleObject.ProgId == "PowerPoint.Show.12":
stream.Save(outputFile_px)
elif oleObject.ProgId == "Package":
stream.Save(outputFile_zip)
stream.Dispose()
# 释放资源
presentation.Dispose()
通过修改 PowerPoint 中 OLE 对象的数据,您可以使演示更具活力、实用性和个性化,为观众呈现更加精彩和引人入胜的演示内容。具体实现步骤如下:
from spire.presentation.common import *
from spire.presentation import *
# 创建一个新的 PowerPoint 演示文稿对象
presentation = Presentation()
# 加载示例文档
presentation.LoadFromFile("示例文档.pptx")
# 遍历每一页的Shape
for slide in presentation.Slides:
for shape in slide.Shapes:
# 检查Shape是否为 OLE 对象
if isinstance(shape, IOleObject):
oleObject = shape if isinstance(shape, IOleObject) else None
stream = oleObject.Data
stream2 = Stream()
# 如果是 PowerPoint 文件对象
if oleObject.ProgId == "PowerPoint.Show.12":
# 创建一个新的 PowerPoint 对象
ppt = Presentation()
ppt.LoadFromStream(stream, FileFormat.Auto)
# 向第一页添加嵌入图像
ppt.Slides[0].Shapes.AppendEmbedImageByPath(ShapeType.Rectangle, "Logo.png", RectangleF.FromLTRB(567, 267, 657, 367))
# 保存修改后的 PowerPoint 文件到流中
ppt.SaveToFile(stream2, FileFormat.Pptx2013)
stream2.Position = 0
oleObject.Data = stream2
# 保存修改后的 PowerPoint 演示文稿
presentation.SaveToFile("修改OLE数据.pptx", FileFormat.Pptx2013)
presentation.Dispose()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。

Word 文档(.doc 和 .docx)是企业和个人办公中最常用的文件格式,广泛应用于合同、报告、手册、技术文档和教育资料。在 C# 开发中,开发者经常需要以编程方式读取 Word 文件内容,例如提取文本用于搜索和分析、获取表格和图片进行数据处理,或抽取批注和文档元数据用于审计、归档和报表生成。
本文将详细介绍如何使用 C# 与 Spire.Doc for .NET 库读取 Word 文档,涵盖文档加载、全文本提取、段落及格式信息获取、表格与图片提取、批注和文档元数据访问,以及页眉和页脚内容读取的完整方法和示例。通过这些方法,开发者可以高效实现 Word 文档内容的自动化处理、数据抽取和信息管理。
在开始之前,需要准备好 C# 开发环境(如 Visual Studio)并安装 Spire.Doc for .NET 库。Spire.Doc for .NET是一个功能完善的 .NET Word文档处理库,支持读取和操作 Word .doc 与 .docx 文件,且无需依赖 Microsoft Word。
使用 Spire.Doc,开发者可以:
安装方法
在 Package Manager Console 中运行以下命令即可从 NuGet 安装 Spire.Doc:
PM> Install-Package Spire.Doc
安装完成后,即可在 C# 项目中调用 Spire.Doc 的 API 解析 Word 文档内容。
要读取 Word 文档,首先需要将文件加载到 Document 对象中。Document 类是 Spire.Doc 的核心类,表示一个 Word 文档实例,提供访问和操作文档内容的入口。
以下示例展示了如何使用 C# 加载一个 .docx 或 .doc 文档:
using Spire.Doc;
using Spire.Doc.Documents;
using System;
namespace LoadWordExample
{
class Program
{
static void Main(string[] args)
{
// 指定 Word 文档的路径
string filePath = @"C:\Documents\示例.docx";
// 创建 Document 对象
using (Document document = new Document())
{
// 加载 Word .docx 或 .doc 文档
document.LoadFromFile(filePath);
}
}
}
}
将 Word 文档加载到 Document 对象后,可以访问文档中的各种内容。根据需求,可以选择全文提取、按段落提取、表格或图片提取等不同方式,以满足全文搜索、版式保留、数据分析或文档迁移的需求。
当目标是全文索引或快速分析时,可以使用 Document.GetText() 方法将 Word 文档中所有文本内容提取为纯文本。此方法保留段落换行,但忽略格式信息、图片或表格结构,适合全文搜索、简单文本统计和分析等场景。
以下示例展示了如何提取Word文档中的所有文本内容:
using (StreamWriter writer = new StreamWriter("提取文本.txt", false, Encoding.UTF8))
{
// 从文档中获取所有文本
string allText = document.GetText();
// 将完整的文本写入到文件中
writer.Write(allText);
}

按段落读取可以获取文档的逻辑结构:段落文本、段落样式(如对齐方式、行间距/段间距)、甚至文本级别的字体样式。适用于需要保留版式信息的场景,例如将 Word 内容导出到自定义渲染器、或根据段落格式作条件处理(如只抽取标题段落)。
以下示例展示了如何按段落读取Word文档并获取段落格式如对齐方式、段后间距:
using (StreamWriter writer = new StreamWriter("按段落读取.txt", false, Encoding.UTF8))
{
// 遍历文档中的所有节
foreach (Section section in document.Sections)
{
// 遍历该节中的所有段落
foreach (Paragraph paragraph in section.Paragraphs)
{
// 获取段落的对齐方式
HorizontalAlignment alignment = paragraph.Format.HorizontalAlignment;
// 获取段落后的间距
float afterSpacing = paragraph.Format.AfterSpacing;
// 将段落的格式信息和文本写入文件
writer.WriteLine($"[对齐方式: {alignment}, 段后间距: {afterSpacing}]");
writer.WriteLine(paragraph.Text);
writer.WriteLine(); // 在段落之间添加一个空行
}
}
}
Word 文档中的图片通常用于图表、说明或者合同签章。在Spire.Doc中,图片由DocPicture对象表示。开发者可查找文档中的 DocPicture 对象,并将其保存为独立的图片文件(支持JPG、PNG等多种格式)。
// 如果文件夹不存在,则创建文件夹
string imageFolder = "提取图片";
if (!Directory.Exists(imageFolder))
Directory.CreateDirectory(imageFolder);
int imageIndex = 0;
// 遍历节和段落以查找图片
foreach (Section section in document.Sections)
{
foreach (Paragraph paragraph in section.Paragraphs)
{
foreach (DocumentObject obj in paragraph.ChildObjects)
{
if (obj is DocPicture picture)
{
// 将每张图片保存为单独的 PNG 文件
string fileName = Path.Combine(imageFolder, $"图_{imageIndex}.png");
picture.Image.Save(fileName, System.Drawing.Imaging.ImageFormat.Png);
imageIndex++;
}
}
}
}

Word 文档中的表格通常用于展示高度结构化的数据,例如财务报表、统计分析表或问卷调查结果。在程序化处理文档时,开发者常常需要提取表格中的文本内容,用于生成报表、导入数据库或进行进一步的数据分析。
使用 Spire.Doc,开发者可以在处理每个表格时,按行(TableRow)和单元格(TableCell)依次提取表格内容,从而保留表格的逻辑结构,便于将数据导出为 CSV、文本或其他可分析的格式,实现对 Word 表格内容的高效处理与自动化应用。
下面的示例展示了如何提取 Word 文档中的所有表格数据,并将每个表格数据保存为独立的文本文件:
// 创建用于存放表格的文件夹
string tableDir = "表格";
if (!Directory.Exists(tableDir))
Directory.CreateDirectory(tableDir);
// 遍历每个节
for (int sectionIndex = 0; sectionIndex < document.Sections.Count; sectionIndex++)
{
Section section = document.Sections[sectionIndex];
TableCollection tables = section.Tables;
// 遍历该节中的所有表格
for (int tableIndex = 0; tableIndex < tables.Count; tableIndex++)
{
ITable table = tables[tableIndex];
string fileName = Path.Combine(tableDir, $"Section{sectionIndex + 1}_Table{tableIndex + 1}.txt");
using (StreamWriter writer = new StreamWriter(fileName, false, Encoding.UTF8))
{
// 遍历每一行
for (int rowIndex = 0; rowIndex < table.Rows.Count; rowIndex++)
{
TableRow row = table.Rows[rowIndex];
// 遍历每个单元格
for (int cellIndex = 0; cellIndex < row.Cells.Count; cellIndex++)
{
TableCell cell = row.Cells[cellIndex];
// 遍历单元格中的每个段落
for (int paraIndex = 0; paraIndex < cell.Paragraphs.Count; paraIndex++)
{
writer.Write(cell.Paragraphs[paraIndex].Text.Trim() + " ");
}
// 单元格之间添加制表符
if (cellIndex < row.Cells.Count - 1) writer.Write("\t");
}
// 每行结束后换行
writer.WriteLine();
}
}
}
}

如需将Word表格数据导出为Excel表格进行进一步数据分析,可参考教程:C# 从 Word 文档中提取表格。
批注(评论)是团队协作时记录反馈的常用方式。程序化读取批注便于审计、迁移评论或生成审阅报告。
Document 对象提供 Comments 集合,可访问 Word 文档中的所有批注。每条批注可能包含多个段落,开发者可以提取每个段落的文本进行进一步处理或保存到文件中。
using (StreamWriter writer = new StreamWriter("批注.txt", false, Encoding.UTF8))
{
// 遍历文档中的所有批注
foreach (Comment comment in document.Comments)
{
// 遍历批注中的每个段落
foreach (Paragraph p in comment.Body.Paragraphs)
{
writer.WriteLine(p.Text);
}
// 添加空行以区分不同的批注
writer.WriteLine();
}
}
文档元数据(如标题、作者、主题、关键字)对文档管理、检索与归档非常重要。这些元数据可以通过 Document 对象的 BuiltinDocumentProperties 属性访问。
using (StreamWriter writer = new StreamWriter("文档元数据.txt", false, Encoding.UTF8))
{
// 将内置文档属性写入文件
writer.WriteLine("标题: " + document.BuiltinDocumentProperties.Title);
writer.WriteLine("作者: " + document.BuiltinDocumentProperties.Author);
writer.WriteLine("主题: " + document.BuiltinDocumentProperties.Subject);
}
页眉和页脚是 Word 文档中常见的版式元素,通常用于显示文档标题、章节名称、作者信息或版权声明。在自动化处理文档时,读取页眉和页脚的内容可以帮助开发者生成完整的文档摘要、制作自定义报告,或者在文档迁移与分析中保留页面结构信息。
使用 Spire.Doc,开发者可以通过 Section.HeadersFooters.Header 或 Section.HeadersFooters.Footer 属性访问每个节的页眉或页脚,然后遍历其中的段落来获取文本内容。
下面的示例展示了如何将文档中每个节的页眉和页脚内容提取到文本文件中:
using (StreamWriter writer = new StreamWriter("页眉页脚.txt", false, Encoding.UTF8))
{
// 遍历文档中的所有节
foreach (Section section in document.Sections)
{
// 写入页眉段落
foreach (Paragraph headerParagraph in section.HeadersFooters.Header.Paragraphs)
{
writer.WriteLine("页眉: " + headerParagraph.Text);
}
// 写入页脚段落
foreach (Paragraph footerParagraph in section.HeadersFooters.Footer.Paragraphs)
{
writer.WriteLine("页脚: " + footerParagraph.Text);
}
}
}
在实际项目中,读取 Word 文档还需要注意以下几点:
通过本文的方法,开发者可以使用 C# 高效读取 Word 文档中的各种内容,包括文本、段落及格式信息、表格、图片、批注、文档元数据,以及页眉页脚。借助 Spire.Doc for .NET,内容提取不仅高效,还能保证数据完整与准确。无论是进行文档分析、生成报表、迁移数据,还是在项目中自动化处理 Word 文件,这些方法都能为开发者提供可靠且实用的解决方案。
Q1: 读取 Word 文档是否需要安装 Microsoft Word?
不需要。Spire.Doc 是独立库,运行环境无需安装 Office。
Q2: 是否可以只提取 Word 文档中的部分内容?
可以。开发者可按段落、表格、节甚至页面进行选择性读取。
Q3: 提取图片时是否会损失质量?
不会。Spire.Doc 提供原始图片数据提取,保证清晰度不受影响。
Q4: 是否支持读取加密的 Word 文档?
支持。只需在加载文档时传入密码即可正常读取内容。
Spire.PDF for Java 10.2.6 已发布。本次更新修复了一些已知问题,如验证签名的结果不正确的问题。详情请阅读以下内容。
问题修复:
Spire.Doc for C++ 12.2.1 已发布。该版本支持通过固定布局获取页的内容。详情请阅读以下内容。
新功能:
// 指定文件路径
wstring input_path = DATAPATH;
wstring inputFile = input_path + L"in.docx";
wstring output_path = OUTPUTPATH;
wstring outputFile = output_path + L"out.txt";
// 创建一个新的 Document 实例
intrusive_ptr<Document> document = new Document();
// 从指定文件加载文档
document->LoadFromFile(inputFile.c_str(), FileFormat::Docx);
intrusive_ptr<FixedLayoutDocument> layoutDoc = new FixedLayoutDocument(document);
wstring result;
// 使用加载的文档创建一个FixedLayoutDocument对象
intrusive_ptr<FixedLayoutLine> line = layoutDoc->GetPages()->GetItem(0)->GetColumns()->GetItem(0)->GetLines()->GetItem(0);
result.append(L"Line: ");
result.append(line->GetText());
result.append(L"\n");
// 检索与该行关联的原始段落
intrusive_ptr<Paragraph> para = line->GetParagraph();
result.append(L"Paragraph text: ");
result.append(para->GetText());
result.append(L"\n");
// 以纯文本格式检索第一页上出现的所有文本(包括页眉和页脚)
wstring pageText = layoutDoc->GetPages()->GetItem(0)->GetText();
result.append(pageText);
result.append(L"\n");
// 循环遍历文档中的每一页并打印每页上出现的行数
for (int i = 0; i < layoutDoc->GetPages()->GetCount(); i++)
{
intrusive_ptr<FixedLayoutPage> page = layoutDoc->GetPages()->GetItem(i);
intrusive_ptr<LayoutCollection> lines = page->GetChildEntities(LayoutElementType::Line, true);
result.append(L"Page ");
result.append(std::to_wstring(page->GetPageIndex()));
result.append(L" has ");
result.append(std::to_wstring(lines->GetCount()));
result.append(L" lines.");
result.append(L"\n");
}
// 对第一段的布局实体执行反向查找
result.append(L"\n");
result.append(L"The lines of the first paragraph:");
result.append(L"\n");
intrusive_ptr<Paragraph> para2 = (Object::Dynamic_cast<Section>(document->GetFirstChild()))->GetBody()->GetParagraphs()->GetItemInParagraphCollection(0);
intrusive_ptr<LayoutCollection> paragraphLines = layoutDoc->GetLayoutEntitiesOfNode(para2);
for (int i = 0; i < paragraphLines->GetCount(); i++)
{
intrusive_ptr<FixedLayoutLine> paragraphLine = Object::Dynamic_cast<FixedLayoutLine>(paragraphLines->GetItem(i));
result.append(paragraphLine->GetText());
result.append(L"\n");
result.append(paragraphLine->GetRectangle()->ToString());
result.append(L"\n");
result.append(L"\n");
}
// 将提取的文本写入文件
std::wofstream write(outputFile);
auto LocUtf8 = locale(locale(""), new std::codecvt_utf8<wchar_t>);
write.imbue(LocUtf8);
write << result;
write.close();
// 处理文档资源
document->Dispose();Spire.Doc for Python 12.2.1 已发布。本次更新新增支持通过固定布局获取页的内容。详情请阅读以下内容。
新功能:
def WriteAllText(fpath:str,content:str):
with open(fpath,'w',encoding="utf-8") as fp:
fp.write(content)
# 指定文件路径
inputFile = "./Data/Sample.docx"
outputFile = "output.txt"
# 创建一个新的 Document 实例
doc = Document()
# 从指定文件加载文档
doc.LoadFromFile(inputFile, FileFormat.Docx)
# 使用加载的文档创建一个 FixedLayoutDocument 对象
layoutDoc = FixedLayoutDocument(doc)
result = ''
# 获取第一页第一列的第一行
line = layoutDoc.Pages[0].Columns[0].Lines[0]
result += "行: "
result += line.Text
result += "\n"
# 获取与该行关联的原始段落
para = line.Paragraph
result += "段落文本: "
result += para.Text
result += "\n"
# 获取以纯文本格式显示在第一页上的所有文本(包括页眉和页脚)。
pageText = layoutDoc.Pages[0].Text
result += pageText
result += "\n"
# 遍历文档中的每一页,并打印每页上出现的行数。
pages = layoutDoc.Pages
for i in range(pages.Count):
page = pages[i]
lines = page.GetChildEntities(LayoutElementType.Line, True)
result += "第 "
result += str(page.PageIndex)
result += " 页有 "
result += str(lines.Count)
result += " 行。"
result += "\n"
# 对第一个段落执行反向查找布局实体
result += "\n"
result += "第一个段落的行:"
result += "\n"
tempChild = doc.FirstChild
section = Section(tempChild)
para = section.Body.Paragraphs[0]
paragraphLines = layoutDoc.GetLayoutEntitiesOfNode(para)
for i in range(paragraphLines.Count):
tempLine = paragraphLines[i]
paragraphLine = FixedLayoutLine(tempLine)
result += (paragraphLine.Text).strip()
result += "\n"
result += paragraphLine.Rectangle.ToString()
result += "\n"
result += "\n"
# 将提取的文本写入文件
WriteAllText(outputFile, result)
# 释放文档资源
doc.Dispose()https://www.e-iceblue.cn/Downloads/Spire-Presentation-Python.html
Spire.XLS for Java 14.2.4已发布。该版本新增支持保存Kingdraw绘制的OLE对象为图片。此外,一些已知问题也在该版本中被成功修复,如将Excel转换为PDF后内容不正确的问题。详情请阅读以下内容。
新功能:
com.spire.xls.Workbook workbook = new com.spire.xls.Workbook();
workbook.loadFromFile("data.xlsx");
Worksheet sheet = workbook.getWorksheets().get(0);
Object o = sheet.getCellRange("C2").getFormulaValue();
if (sheet.hasOleObjects()) {
for (int i = 0; i < sheet.getOleObjects().size(); i++) {
IOleObject oleObject = sheet.getOleObjects().get(i);
OleObjectType oleObjectType = sheet.getOleObjects().get(i).getObjectType();
byte[] picUrl = null;
switch (oleObjectType) {
case Emf:
picUrl = oleObject.getOleData();;
break;
}
if (picUrl != null) {
byteArrayToFile(picUrl, "out.png");
break;
}
}
}
}
public static void byteArrayToFile(byte[] datas, String destPath) {
File dest = new File(destPath);
try (InputStream is = new ByteArrayInputStream(datas);
OutputStream os = new BufferedOutputStream(new FileOutputStream(dest, false));) {
byte[] flush = new byte[1024];
int len = -1;
while ((len = is.read(flush)) != -1) {
os.write(flush, 0, len);
}
os.flush();
} catch (IOException e) {
e.printStackTrace();
}
}问题修复:
在 PDF 文档中,除了使用文本和图像构成文档基本内容外,用户还可以在其中嵌入各种类型的附件,如文档、图像、音频文件或其他多媒体元素。许多 PDF 文档会嵌入附件作为补充资料,如报告、电子表格或法律文件。将这些文档中的附件提取出来,可以方便用户对附件进行检索、修改和分享等操作。本文将演示如何使用 Spire.PDF for Python 通过 Python 程序从 PDF 文档中提取附件。
本教程需要用到 Spire.PDF for Python 和 plum-dispatch v1.7.4。可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.PDF如果您不确定如何安装,请参考:如何在 Windows 中安装 Spire.PDF for Python
PDF 文件中有两类附件:文档级附件和注释级附件。下面的表格说明了这两类附件之间的差异以及它们在 Spire.PDF 中的表示方式。
| 附件类型 | 表示方式 | 定义 |
| 文档附件 | PdfAttachment 类 | 以文档级添加的 PDF 附件不会显示在 PDF 页面上,但可以在 PDF 阅读器的“附件”面板中查看。 |
| 注释附件 | PdfAnnotationAttachment 类 | 作为注释附加的文件可以在页面上或“附件”面板中找到。注释附件在页面上显示为一个纸夹图标;阅读文档时可以双击该图标打开文件。 |
使用 PdfDocument.Attachments 属性可以获取 PDF 文档中的文档附件。每个附件都有一个 PdfAttachment.FileName 属性,提供指定附件的文件名(包括文件扩展名)。此外,PdfAttachment.Data 属性能帮助开发者访问附件的数据,PdfAttachment.Data.Save() 方法将附件保存到指定文件夹。
使用 Python 从 PDF 中提取文档附件的操作步骤如下:
from spire.pdf import *
from spire.pdf.common import *
# 创建PdfDocument类的对象
doc = PdfDocument()
# 载入PDF文件
doc.LoadFromFile("示例.pdf")
# 获取文档中的文档附件集合
collection = doc.Attachments
# 遍历附件集合
if collection.Count > 0:
for i in range(collection.Count):
# 获取指定附件
attachment = collection.get_Item(i)
# 获取附件的文件名和文件数据
fileName= attachment.FileName
data = attachment.Data
# 保存附件到指定文件夹
data.Save("output/文档附件/" + fileName)
doc.Close()
注释附件是基于页面的元素。要从特定页面获取注释附件,可以先通过 PdfPageBase.AnnotationsWidget 属性获取页面注释,然后判断注释是否是附件注释,最后将附件注释中的附件保存到指定的文件夹即可。
以下是使用 Python 从 PDF 中提取注释附件的操作步骤:
from spire.pdf import *
from spire.pdf.common import *
# 创建PdfDocument类的对象
doc = PdfDocument()
# 载入PDF文件
doc.LoadFromFile("示例1.pdf")
# 遍历文档页面
for i in range(doc.Pages.Count):
# 获取指定页面
page = doc.Pages.get_Item(i)
# 获取页面上的所有注释
annotationCollection = page.AnnotationsWidget
# 判断页面是否包含注释
if annotationCollection.Count > 0:
# 遍历页面中的注释
for j in range(annotationCollection.Count):
# 获取指定注释
annotation = annotationCollection.get_Item(j)
# 判断指定注释是否为附件注释
if isinstance(annotation, PdfAttachmentAnnotationWidget):
# 获取附件的文件名和文件数据
fileName = annotation.FileName
byteData = annotation.Data
streamMs = Stream(byteData)
# 将附件保存到指定文件夹
streamMs.Save("output/注释附件/" + fileName)
doc.Close()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
OLE 对象是一种可以插入到文档中的来自其他应用程序的文件或数据。这些对象可以是 Excel 表格、幻灯片、PDF文档、图像、音频、视频等。通过将 OLE 对象插入到 Word 文档中,您可以在文档中直接打开、编辑和操作这些对象,而无需打开原始文件或应用程序。在这篇文章中,我们将介绍如何使用 Python 和 Spire.Doc for Python 在 Word 文档中插入和提取 OLE 对象。
本教程需要 Spire.Doc for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.Doc如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.Doc for Python
Spire.Doc for Python 提供了 Paragraph.AppendOleObject(pathToFile:str, olePicture:DocPicture, type:OleObjectType) 方法,用于在 Word 文档中嵌入 OLE 对象。具体步骤如下:
以下代码示例展示了如何使用 Spire.Doc for Python 在 Word 文档中嵌入 Excel 电子表格、PDF 文件和 PowerPoint 演示文稿:
from spire.doc import *
from spire.doc.common import *
# 创建Document类的对象
doc = Document()
# 加载Word文档
doc.LoadFromFile("示例.docx")
# 获取第一个节
section = doc.Sections.get_Item(0)
# 向节添加一个段落
para1 = section.AddParagraph()
para1.AppendText("Excel文件: ")
# 加载一张图片作为OLE对象的图标
picture1 = DocPicture(doc)
picture1.LoadImage("https://cdn.e-iceblue.cn/Excel图标.png")
picture1.Width = 50
picture1.Height = 50
# 向段落附加一个OLE对象(Excel文件)
para1.AppendOleObject("预算.xlsx", picture1, OleObjectType.ExcelWorksheet)
# 向节添加一个段落
para2 = section.AddParagraph()
para2.AppendText("PDF文件: ")
# 加载将用作OLE对象封面图像的图像
picture2 = DocPicture(doc)
picture2.LoadImage("https://cdn.e-iceblue.cn/PDF图标.png")
picture2.Width = 50
picture2.Height = 50
# 向段落附加一个OLE对象(PDF文件)
para2.AppendOleObject("Report.pdf", picture2, OleObjectType.AdobeAcrobatDocument)
# 向节添加一个段落
para3 = section.AddParagraph()
para3.AppendText("PPT文件: ")
# 加载将用作OLE对象封面图像的图像
picture3 = DocPicture(doc)
picture3.LoadImage("https://cdn.e-iceblue.cn/PPT-Icon.png")
picture3.Width = 50
picture3.Height = 50
# 向段落附加一个OLE对象(PowerPoint演示文稿)
para3.AppendOleObject("Input.pptx", picture3, OleObjectType.PowerPointPresentation)
doc.SaveToFile("插入OLE.docx", FileFormat.Docx2013)
doc.Close()
要从 Word 文档中提取 OLE 对象,首先需要识别出文档中的 OLE 对象。识别后,可以判断每个 OLE 对象的文件格式,最后将每个 OLE 对象的数据以其原生文件格式保存到文件中。具体步骤如下:
以下代码示例展示了如何使用 Spire.Doc for Python 从 Word 文档中提取嵌入的 Excel 表格、PDF 文件和 PowerPoint 演示文稿:
from spire.doc import *
from spire.doc.common import *
# 创建Document类的对象
doc = Document()
# 加载Word文档
doc.LoadFromFile("插入OLE.docx")
i = 1
# 遍历Word文档的所有节
for k in range(doc.Sections.Count):
sec = doc.Sections.get_Item(k)
# 遍历每个节的所有子对象
for j in range(sec.Body.ChildObjects.Count):
obj = sec.Body.ChildObjects.get_Item(j)
# 检查子对象是否为段落
if isinstance(obj, Paragraph):
par = obj if isinstance(obj, Paragraph) else None
# 遍历段落中的子对象
for m in range(par.ChildObjects.Count):
o = par.ChildObjects.get_Item(m)
# 检查子对象是否为OLE对象
if o.DocumentObjectType == DocumentObjectType.OleObject:
ole = o if isinstance(o, DocOleObject) else None
s = ole.ObjectType
# 检查OLE对象是否为PDF文件
if s.startswith("AcroExch.Document"):
ext = ".pdf"
# 检查OLE对象是否为Excel表格
elif s.startswith("Excel.Sheet"):
ext = ".xlsx"
# 检查OLE对象是否为PowerPoint演示文稿
elif s.startswith("PowerPoint.Show"):
ext = ".pptx"
else:
continue
# 将OLE的数据以其原生格式写入文件
with open(f"Output/OLE{i}{ext}", "wb") as file:
file.write(ole.NativeData)
i += 1
doc.Close()
如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。