在文档自动化场景中,精准的文本对齐对于创建专业、易读且视觉吸引力强的文档至关重要。对于开发者而言,在生成报告、起草信函或设计单据时,能够通过 Python 设置文本对齐样式可以实现批量生成规范文档。本文将介绍如何通过 Spire.Doc for Python(一款可灵活控制 Word 文档格式的库)在 Python 中实现文本对齐的操作。
在编写代码前,我们先来明确为什么 Spire.Doc 是文本对齐任务的首选:
Spire.Doc 通过 HorizontalAlignment 枚举来定义文本对齐方式,常用取值包括:
HorizontalAlignment.Left:文本左对齐(默认对齐方式)。HorizontalAlignment.Right:文本右对齐。HorizontalAlignment.Center:文本在左右边距间水平居中。HorizontalAlignment.Justify:自动调整文本间距,使左右两边均与边距对齐。HorizontalAlignment.Distribute:均匀分配字符间距(字母间加空格)和单词间距,填满整行。下文将具体介绍如何使用 Python 在 Word 中通过编程设置段落对齐方式(左对齐、右对齐、居中对齐、两端对齐、分散对齐)。
以下步骤将生成一个包含 5 个段落的 Word 文档,每个段落对应一种对齐样式。
打开终端/命令提示符,运行以下命令安装最新版本:
pip install Spire.Doc
从 Spire.Doc 中导入核心类,这些模块用于创建文档、节、段落及配置格式:
from spire.doc import *
from spire.doc.common import *
初始化一个代表空 Word 文件的 Document 实例:
# 创建 Document 实例
doc = Document()
Word 文档通过节组织内容(每个节可独立设置边距、页面大小等)。这里将添加一个节用于容纳段落:
# 向文档添加节
section = doc.AddSection()
一个节可包含多个段落,每个段落的对齐方式通过 HorizontalAlignment 枚举控制。下面创建 5 个段落,分别对应五种对齐类型。
左对齐是多数文本的默认样式(文本向左边距对齐):
# 左对齐文本(默认样式)
paragraph1 = section.AddParagraph()
paragraph1.AppendText("该段落为左对齐。")
paragraph1.Format.HorizontalAlignment = HorizontalAlignment.Left
右对齐适用于日期、签名或页码(文本向右边距对齐):
# 右对齐文本
paragraph2 = section.AddParagraph()
paragraph2.AppendText("该段落为右对齐。")
paragraph2.Format.HorizontalAlignment = HorizontalAlignment.Right
居中对齐适用于标题(文本在左右边距间居中),实现代码如下:
# 居中对齐文本
paragraph3 = section.AddParagraph()
paragraph3.AppendText("该段落为居中对齐。")
paragraph3.Format.HorizontalAlignment = HorizontalAlignment.Center
两端对齐使文本左右两边均与边距对齐(通过调整单词间距保持一致性),适用于论文、报告等正式文档:
# 两端对齐文本
paragraph4 = section.AddParagraph()
paragraph4.AppendText("该段落为两端对齐。")
paragraph4.Format.HorizontalAlignment = HorizontalAlignment.Justify
注意:两端对齐在长文本中效果更明显,短句可能难以体现间距调整。
分散对齐与两端对齐类似,但会均匀分布单行文本(如间距不均的单词或短语):
# 分散对齐文本
paragraph5 = section.AddParagraph()
paragraph5.AppendText("该段落为分散对齐。")
paragraph5.Format.HorizontalAlignment = HorizontalAlignment.Distribute
最后,将文档保存到指定路径,并关闭 Document 实例以释放资源:
# 保存文档
doc.SaveToFile("对齐文本.docx", FileFormat.Docx2016)
# 关闭文档释放内存
doc.Close()
输出效果:

Spire.Doc for Python 还提供了接口用于在 Word 中对齐表格或对齐单元格中的文本。
答:Spire.Doc 提供 免费版,但有功能限制。如需完整功能,可 申请 30 天试用许可证。
答:可以。Spire.Doc 支持加载现有文档并修改特定段落的对齐方式,示例如下:
from spire.doc import *
# 加载现有文档
doc = Document()
doc.LoadFromFile("ExistingDocument.docx")
# 获取第一个节和第一个段落
section = doc.Sections[0]
paragraph = section.Paragraphs[0]
# 修改为居中对齐
paragraph.Format.HorizontalAlignment = HorizontalAlignment.Center
# 保存修改后的文档
doc.SaveToFile("UpdatedDocument.docx", FileFormat.Docx2016)
doc.Close()
答:不能。在 Word 中,文本对齐是段落级属性,同一段落内的所有文本必须采用相同的对齐方式。若需在同一行混合对齐,可使用无边框表格实现。
答:支持!Spire.Doc 可将对齐方式与字体、行间距、项目符号等格式结合使用,具体可参考 字体设置、段落或行间距设置 等教程。
借助 Python 和 Spire.Doc 实现 Word 文本对齐自动化,既能节省大量手动调整时间,减少人为误差,又能保证文档格式的一致性。本文提供的代码示例为左对齐、右对齐、居中对齐、两端对齐和分散对齐提供了清晰模板,只需修改文本或补充格式规则,即可快速适配实际需求。建议尝试不同对齐方式的组合,并探索 Spire.Doc for Python 的在线教程,以解锁更多格式设置的可能性。
对于数据分析和报告而言,视觉美学在有效呈现信息方面发挥着重要作用。在使用 Excel 工作表时,设置背景颜色和图像的功能可以增强数据的整体可读性和影响力。利用 Python 的强大功能,开发人员可以毫不费力地操作 Excel 文件并自定义工作表的外观。本文将演示如何使用 Spire.XLS for Python 通过 Python 程序为 Excel 工作表设置背景颜色和图像。
本教程需要 Spire.XLS for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.XLS
如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.XLS for Python
使用 Spire.XLS for Python,开发人员可以通过 CellRange.Style.Color 属性为指定的单元格区域设置背景色。为工作表中使用的单元格区域设置背景色的详细步骤如下:
from spire.xls import *
from spire.xls.common import *
# 创建一个Workbook对象
wb = Workbook()
# 从输入文件加载Workbook对象
wb.LoadFromFile("输入文档.xlsx")
# 获取Workbook中的第一个工作表
sheet = wb.Worksheets.get_Item(0)
# 获取工作表中已使用的单元格范围
usedRange = sheet.AllocatedRange
# 设置单元格范围的背景颜色为淡绿色
usedRange.Style.Color = Color.FromRgb(144, 238, 144)
# 将修改后的Workbook保存到名为"Excel背景颜色.xlsx"的文件中,指定文件格式为Excel 2016版本
wb.SaveToFile("Excel背景颜色.xlsx", FileFormat.Version2016)
# 释放Workbook对象
wb.Dispose()

为 Excel 工作表设置背景图像可以通过 PageSetup 类来完成。使用 Worksheet.PageSetup.BackgroundImage 属性,开发人员可以为整个工作表设置背景图像。具体步骤如下:
from spire.xls import *
from spire.xls.common import *
# 创建一个Workbook对象
wb = Workbook()
# 从输入文件加载Workbook对象
wb.LoadFromFile("输入文档.xlsx")
# 获取Workbook中的第一个工作表
sheet = wb.Worksheets.get_Item(0)
# 从文件加载背景图片
image = Stream("背景图片.png")
# 将背景图片设置为工作表的背景图像
sheet.PageSetup.BackgoundImage = image
# 将修改后的Workbook保存到名为"Excel背景图片.xlsx"的文件中,指定文件格式为Excel 2016版本
wb.SaveToFile("Excel背景图片.xlsx", FileFormat.Version2016)
# 释放Workbook对象
wb.Dispose()

如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
Excel 的自动筛选(AutoFilter) 功能是数据处理的高效工具,可根据自定义条件快速筛选工作表数据 — 应用筛选后,仅显示符合条件的行,其余数据自动隐藏,极大简化了数据筛选与分析流程。
然而,若长期保留激活状态的自动筛选,可能导致数据误读(如将筛选后数据视为完整数据集)或格式混乱。因此,掌握 “添加” 与 “删除” 自动筛选的操作同样重要。本文将基于 Spire.XLS for Python 库,详细讲解如何在 Python 中实现 Excel 自动筛选器添加和删除操作,附完整代码与场景说明。
目录:
Spire.XLS for Python 是一款轻量且功能全面的 Excel 处理库,支持自动筛选、公式计算、图表生成等近百种 Excel 操作,无需依赖 Microsoft Excel 环境即可运行。
若要安装该 Python 库,请打开终端或命令提示符,运行以下命令:
pip install Spire.XLS
pip 工具会自动从 Python 包索引(PyPI)搜索 Spire.XLS 库的最新版本,然后下载并安装该库及其所有必需的依赖项。
Excel 自动筛选可应用于指定单元格区域(如 A1:C1)或整列(如 A 列)。以下是使用到的核心属性:
Worksheet.AutoFilters:获取工作表中的自动筛选器集合,返回一个 AutoFiltersCollection 对象。AutoFiltersCollection.Range:指定需要筛选的单元格区域。Python 代码示例:
from spire.xls import *
from spire.xls.common import *
# 指定输入文件和输出文件名
inputFile = "输入文档.xlsx"
outputFile = "Excel自动筛选器.xlsx"
# 创建 Workbook 实例
workbook = Workbook()
# 加载 Excel 文件
workbook.LoadFromFile(inputFile)
# 获取第一个工作表
sheet = workbook.Worksheets[0]
# 在工作表中创建自动筛选,并指定要筛选的区域为第一行的A到C列
sheet.AutoFilters.Range = sheet.Range["A1:C1"]
# 保存结果文件
workbook.SaveToFile(outputFile, ExcelVersion.Version2016)
# 释放资源
workbook.Dispose()
结果: 打开输出文件后,会发现A1、B1、C1 单元格右侧出现下拉箭头,点击箭头可展开筛选选项。

Spire.XLS 的 AutoFiltersCollection 类提供了多种筛选方法,覆盖文本、日期、空白值、颜色等常见场景,满足不同数据筛选需求。
| 筛选类型 | 详细说明 |
|---|---|
| 文本筛选 | AddFilter() 方法:筛选包含指定文本内容的单元格。 |
| 日期筛选 | AddDateFilter() 方法:筛选与指定年/月/日等关联的日期。 |
| 空白/非空白单元格筛选 |
|
| 按颜色筛选 |
|
| 自定义筛选 | CustomFilter() 方法:按自定义条件筛选。 |
注意:调用上述方法后,需额外执行 Worksheet.AutoFilters.Filter() 才能生效(触发筛选操作)。
若基础筛选无法满足需求,可通过 CustomFilter(column: FilterColumn, operatorType: FilterOperatorType, criteria: Object) 实现自定义条件筛选。以下示例为过滤包含特定文本的数据。
Python 代码示例:
from spire.xls import *
from spire.xls.common import *
# 指定输入文件和输出文件名
inputFile = "输入文档.xlsx"
outputFile = "自定义自动筛选器.xlsx"
# 创建 Workbook 实例
workbook = Workbook()
# 加载 Excel 文件
workbook.LoadFromFile(inputFile)
# 获取第一个工作表
sheet = workbook.Worksheets[0]
# 设置工作表的自动过滤范围为第二列的前6行
sheet.AutoFilters.Range = sheet.Range["B1:B6"]
# 获取要筛选的列
filtercolumn = sheet.AutoFilters[0]
# 设置自定义过滤条件为筛选出包含"鼠标"的内容
strCrt = String("鼠标")
sheet.AutoFilters.CustomFilter(filtercolumn, FilterOperatorType.Equal, strCrt)
# 执行过滤操作
sheet.AutoFilters.Filter()
# 保存结果文件
workbook.SaveToFile(outputFile, ExcelVersion.Version2016)
# 释放资源
workbook.Dispose()
结果:仅显示包含"鼠标"的内容,其他类别数据将被隐藏。

当筛选完成后,建议删除自动筛选(尤其是需分享文件时),避免他人误判数据完整性。Spire.XLS 提供了 AutoFiltersCollection.Clear() 方法一键删除所有筛选器。
核心作用
Python 代码示例:
from spire.xls import *
from spire.xls.common import *
# 指定输入文件和输出文件名
inputFile = "自定义自动筛选器.xlsx"
outputFile = "删除自动筛选器.xlsx"
# 创建 Workbook 实例
workbook = Workbook()
# 加载 Excel 文件
workbook.LoadFromFile(inputFile)
# 获取第一个工作表
sheet = workbook.Worksheets[0]
# 清除工作表的所有自动筛选器
sheet.AutoFilters.Clear()
# 保存结果文件
workbook.SaveToFile(outputFile, ExcelVersion.Version2016)
# 释放资源
workbook.Dispose()
通过 Spire.XLS for Python,可实现 Excel 自动筛选器的 “添加 - 应用 - 删除” 全流程自动化,无需手动操作 Excel,极大提升数据处理效率。
适用场景
如需进一步探索高级功能(如筛选后导出数据),可参考 Spire.XLS for Python 官方文档。
Spire.Cloud 9.4.9 已发布。该版本在 PDF 文档查看器中新增了下边栏的隐藏设置,同时还修复了一个将 MS Word 文档的内容复制粘贴到 Word 编辑器时出现的问题。详情请查看以下内容。
新功能:
问题修复:
修改 PDF 文档以适应各种使用场景的需求是 PDF 文档创建和管理的常见操作。在这些操作中,分割和合并 PDF 页面可以帮助重新组织 PDF 内容,从而方便进行打印、排版等。通过使用 Python 程序,开发人员可以轻松地将一个 PDF 文档中的页面拆分为多个页面,或将多个 PDF 页面合并为一个页面。本文将演示如何使用 Spire.PDF for Python 在 Python 程序中分割和合并 PDF 页面。
本教程需要用到 Spire.PDF for Python 和 plum-dispatch v1.7.4。可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.PDF
如果您不清楚如何安装,请参考:如何在 Windows 中安装 Spire.PDF for Python
通过 Spire.PDF for Python,开发人员可以使用 PdfPageBase.CreateTemplate().Draw(newPage PdfPageBase, PointF) 方法将一个 PDF 页面绘制到新的 PDF 页面上。绘制时,如果当前新页面无法完全容纳原始页面的内容,则会自动创建一个新的页面,并将剩余的内容绘制到该页面上。因此,我们可以创建一个新的 PDF 文档并通过指定页面大小来控制绘制结果,从而实现指定的 PDF 页面水平或垂直分割。
以下是将 PDF 页面垂直分割为两个独立 PDF 页面的操作步骤:
from spire.pdf import *
from spire.pdf.common import *
# 创建PdfDocument类的对象并加载PDF文档
pdf = PdfDocument()
pdf.LoadFromFile("示例1.pdf")
# 获取文档的第一页
page = pdf.Pages.get_Item(0)
# 创建一个新的PDF文档
newPdf = PdfDocument()
# 将新PDF文档的边距设置为0
newPdf.PageSettings.Margins.All = 0.0
# 获取提取的页面的宽度和高度
width = page.Size.Width
height = page.Size.Height
# 将新PDF文档的宽度设置为提取的页面宽度,高度设置为提取页面高度的一半
newPdf.PageSettings.Width = width
newPdf.PageSettings.Height = height / 2
# 向新PDF文档添加一个新页面
newPage = newPdf.Pages.Add()
# 将提取的页面内容绘制到新页面上
page.CreateTemplate().Draw(newPage, PointF(0.0, 0.0))
# 保存新的PDF文档
newPdf.SaveToFile("output/拆分PDF页面.pdf")
pdf.Close()
newPdf.Close()

同样地,开发者也可以通过在同一个 PDF 页面上绘制不同的页面来合并 PDF 页面。需要注意的是,待合并的页面最好是相同宽度或高度,否则需要取最大值以确保正确绘制。
将两个 PDF 页面合并为一个 PDF 页面的详细步骤如下:
from spire.pdf import *
from spire.pdf.common import *
# 创建PdfDocument类的对象并加载PDF文档
pdf = PdfDocument()
pdf.LoadFromFile("示例2.pdf")
# 获取文档的第一页和第二页
page = pdf.Pages.get_Item(0)
page1 = pdf.Pages.get_Item(1)
# 创建一个新的PDF文档
newPdf = PdfDocument()
# 将新PDF文档的边距设置为0
newPdf.PageSettings.Margins.All = 0.0
# 将新文档的页面宽度设置为提取的页面的宽度相同
newPdf.PageSettings.Width = page.Size.Width
# 将新文档的页面高度设置为两个提取页面高度的总和
newPdf.PageSettings.Height = page.Size.Height + page1.Size.Height
# 向新PDF文档添加一个新页面
newPage = newPdf.Pages.Add()
# 将提取的页面的内容绘制到新页面上
page.CreateTemplate().Draw(newPage, PointF(0.0, 0.0))
page1.CreateTemplate().Draw(newPage, PointF(0.0, page.Size.Height))
# 保存新文档
newPdf.SaveToFile("output/合并PDF页面.pdf")
pdf.Close()
newPdf.Close()

如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
Spire.Office 9.4.0 已发布。在该版本中,Spire.Doc 支持加载和操作 Markdown 文档;Spire.PDF 支持获取查找到的文本的字体、字体大小以及字体格式;Spire.Presentation 支持在段落中插入公式以及将 SVG 文件以图片形式嵌入到幻灯片中。此外,大量已知问题也在该版本中成功修复。详情请阅读以下内容。
该版本涵盖了最新版的 Spire.Doc,Spire.PDF,Spire.XLS,Spire.Email,Spire.DocViewer,Spire.PDFViewer,Spire.Presentation,Spire.Spreadsheet,Spire.OfficeViewer,Spire.Barcode,Spire.DataExport。
版本信息如下:
https://www.e-iceblue.cn/Downloads/Spire-Office-NET.html
新功能:
Document doc = new Document();
//加载 .md 文件
doc.LoadFromFile("input.md");
//保存为 .md 文件
//doc.SaveToFile("output.md", Spire.Doc.FileFormat.Markdown);
//保存为 .docx 文件
//doc.SaveToFile("output.docx", Spire.Doc.FileFormat.Docx);
//保存为 .doc 文件
//doc.SaveToFile("output.doc", Spire.Doc.FileFormat.Doc);
//保存为 .pdf 文件
doc.SaveToFile("output.pdf", Spire.Doc.FileFormat.PDF);
doc.Close();
Document doc = new Document();
//加载 .docx 文件
doc.LoadFromFile("input.docx");
//加载 .doc 文件
//doc.LoadFromFile("input.doc");
//保存为 .md 文件
doc.SaveToFile("output.md", Spire.Doc.FileFormat.Markdown);
doc.Close();问题修复:
新功能:
PdfDocument pdf = new PdfDocument();
pdf.LoadFromFile("test.pdf");
PdfTextFindOptions findOptions = new PdfTextFindOptions();
findOptions.Parameter = TextFindParameter.IgnoreCase;
foreach (PdfPageBase page in pdf.Pages)
{
PdfTextFinder finder = new PdfTextFinder(page);
finder.Options = findOptions;
List results = finder.Find("total");
foreach (PdfTextFragment text in results)
{
String font=text.TextStates[0].FontName;
float size = text.TextStates[0].FontSize;
String fontF = text.TextStates[0].FontFamily;
}
} PdfDocument pdf = new PdfDocument();
pdf.LoadFromFile(inputFile);
PdfPageBase page = pdf.Pages[0];
PdfTextFinder finds = new PdfTextFinder(page);
finds.Options.Parameter = TextFindParameter.None;
List<PdfTextFragment> result = finds.Find("hello");
StringBuilder str = new StringBuilder();
foreach (PdfTextFragment find in result)
{
string text = find.LineText;
string FontName = find.TextStates[0].FontName;
float FontSize = find.TextStates[0].FontSize;
string FontFamily = find.TextStates[0].FontFamily;
bool IsBold = find.TextStates[0].IsBold;
bool IsSimulateBold = find.TextStates[0].IsSimulateBold;
bool IsItalic = find.TextStates[0].IsItalic;
Color color = find.TextStates[0].ForegroundColor;
str.AppendLine(text);
str.AppendLine("FontName: " + FontName);
str.AppendLine("FontSize: " + FontSize);
str.AppendLine("FontFamily: " + FontFamily);
str.AppendLine("IsBold: " + IsBold);
str.AppendLine("IsSimulateBold: " + IsSimulateBold);
str.AppendLine("IsItalic: " + IsItalic);
str.AppendLine("color: " + color);
str.AppendLine(" ");
}
PdfTextReplacer ptr = new PdfTextReplacer(page);
ptr.ReplaceAllText("hello", "New");
File.WriteAllText(outputFile_T, str.ToString());
pdf.SaveToFile(outputFile);
pdf.Dispose();PdfDocument pdf = new PdfDocument();
pdf.LoadFromFile(inputFile);
pdf.RegisterProgressNotifier(new CustomProgressNotifier());
pdf.SaveToFile(outputFile, FileFormat.XPS);
pdf.Close();
public class CustomProgressNotifier :IProgressNotifier
{
StringBuilder str = new StringBuilder();
public void Notify(float progress)
{
str.AppendLine(progress + "%");
File.WriteAllText(outputFile_txt, str.ToString());
}
}问题修复:
新功能:
public enum InsertPlaceholderType
{
Content = 0,
VerticalContent = 1,
Text = 2,
VerticalText = 3,
Picture = 4,
Chart = 5,
Table = 6,
SmartArt = 7,
Media = 8,
OnlineImage = 9
}
presentation.Masters[0].Layouts[0].InsertPlaceholder(InsertPlaceholderType.Text, new RectangleF(20, 30, 400, 400));Presentation ppt = new Presentation();
ppt.LoadFromFile(inputFile);
IChart chart = ppt.Slides[0].Shapes[9] as IChart;
ProjectionType type = chart.Series[0].ProjectionType;
chart.Series[0].ProjectionType = ProjectionType.Robinson;
ppt.SaveToFile(outputFile, FileFormat.Pptx2013);
ppt.Dispose();Presentation ppt = new Presentation();
string latexMathCode = "x^{2}+\\sqrt{x^{2}+1=2}";
IAutoShape shape = ppt.Slides[0].Shapes.AppendShape(ShapeType.Rectangle, new RectangleF(30, 100, 400, 200));
shape.TextFrame.Paragraphs.Clear();
TextParagraph p = new TextParagraph();
p.ParagraphProperties.DefaultTextRangeProperties.Fill.FillType = FillFormatType.Solid;
p.ParagraphProperties.DefaultTextRangeProperties.Fill.SolidColor.Color = Color.Black;
shape.TextFrame.Paragraphs.Append(p);
TextRange portionEx = new TextRange("Hello World");
p.TextRanges.Append(portionEx);
p.AppendFromLatexMathCode(latexMathCode);
TextRange portionEx2 = new TextRange("My name is Tom.");
p.TextRanges.Append(portionEx2);
ppt.SaveToFile(outputFile, FileFormat.Auto);
ppt.Dispose();presentation.Slides[0].Shapes.AddFromSVG(inputFile, new RectangleF(40, 40, 200, 200));问题修复:
问题修复:
Spire.Presentation 9.4.5 已发布。该版本新增多项功能,如支持添加占位符,支持在段落中插入公式等。此外还修复了一个在转换 PPTX 文档到 SVG 文档时形状渐变色背景的方向被旋转的问题。详情见下文。
新功能:
public enum InsertPlaceholderType
{
Content = 0,
VerticalContent = 1,
Text = 2,
VerticalText = 3,
Picture = 4,
Chart = 5,
Table = 6,
SmartArt = 7,
Media = 8,
OnlineImage = 9
}
presentation.Masters[0].Layouts[0].InsertPlaceholder(InsertPlaceholderType.Text, new RectangleF(20, 30, 400, 400));Presentation ppt = new Presentation();
ppt.LoadFromFile(inputFile);
IChart chart = ppt.Slides[0].Shapes[9] as IChart;
ProjectionType type = chart.Series[0].ProjectionType;
chart.Series[0].ProjectionType = ProjectionType.Robinson;
ppt.SaveToFile(outputFile, FileFormat.Pptx2013);
ppt.Dispose();Presentation ppt = new Presentation();
string latexMathCode = "x^{2}+\\sqrt{x^{2}+1=2}";
IAutoShape shape = ppt.Slides[0].Shapes.AppendShape(ShapeType.Rectangle, new RectangleF(30, 100, 400, 200));
shape.TextFrame.Paragraphs.Clear();
TextParagraph p = new TextParagraph();
p.ParagraphProperties.DefaultTextRangeProperties.Fill.FillType = FillFormatType.Solid;
p.ParagraphProperties.DefaultTextRangeProperties.Fill.SolidColor.Color = Color.Black;
shape.TextFrame.Paragraphs.Append(p);
TextRange portionEx = new TextRange("Hello World");
p.TextRanges.Append(portionEx);
p.AppendFromLatexMathCode(latexMathCode);
TextRange portionEx2 = new TextRange("My name is Tom.");
p.TextRanges.Append(portionEx2);
ppt.SaveToFile(outputFile, FileFormat.Auto);
ppt.Dispose();presentation.Slides[0].Shapes.AddFromSVG(inputFile, new RectangleF(40, 40, 200, 200));问题修复:
https://www.e-iceblue.cn/Downloads/Spire-Presentation-NET.html
在 Word 文本框中,用户可以自由插入、调整文本内容,实现多样化的文档布局,它不仅支持丰富的文本格式设置,还可与其他文档元素如图片、表格等协同工作。在使用程序操作文档时,我们也经常需要提取出文本框中的内容,将数据保存到数据库或者其他文件中。在这篇文章中,我们将探讨如何使用 Spire.Doc for Python 获取 Word 文本框中的文本、图片、表格。
本教程需要 Spire.Doc for Python 和 plum-dispatch v1.7.4。您可以通过以下 pip 命令将它们轻松安装到 Windows 中。
pip install Spire.Doc
如果您不确定如何安装,请参考此教程: 如何在 Windows 中安装 Spire.Doc for Python
用于测试的 Word 源文档如下图:

使用 Spire.Doc for Python 获取文本框内文本内容的详细步骤如下:
from spire.doc import *
from spire.doc.common import *
# 创建Document对象
document = Document()
# 加载文件
document.LoadFromFile("D:\\schedule\\Data\\提取textbox内容.docx")
# 定义文本输出文件名
outputFile = "ExtractTextFromTextBoxes.txt"
# 如果文档中的文本框个数大于0
if document.TextBoxes.Count > 0:
with open(outputFile, 'w', encoding='utf-8') as sw:
# 遍历文档的章节-段落
for i in range(document.Sections.Count):
section = document.Sections.get_Item(i)
for j in range(section.Paragraphs.Count):
p = section.Paragraphs.get_Item(j)
for k in range(p.ChildObjects.Count):
obj = p.ChildObjects.get_Item(k)
# 判断类型,如果获取到的对象是文本框就转换obj为TextBox对象
if obj.DocumentObjectType == DocumentObjectType.TextBox:
textbox = obj if isinstance(obj, TextBox) else None
for x in range(textbox.ChildObjects.Count):
objt = textbox.ChildObjects.get_Item(x)
# 如果对象是段落,则输出段落文本
if objt.DocumentObjectType == DocumentObjectType.Paragraph:
sw.write((objt if isinstance(objt, Paragraph) else None).Text)
document.Close()

在上面的例子中,为了获取 Textbox 对象,需要遍历整个文档,当我们需要获取的目标对象属于非常基础的对象(比如 DocPicture)时,又需要继续遍历整个 Textbox 对象及其子对象,这样操作会显得非常繁琐。为了简化操作,我们借用 queue.Queue() 方法将 Textbox 压入队列进行遍历,这样只用取出队列中的对象判断是否为图片对象,若不是则获取其子对象继续压入队列。具体请参考下面的步骤:
from spire.doc import *
from spire.doc.common import *
import queue
# 创建Document对象
document = Document()
# 加载文件
document.LoadFromFile("D:\\schedule\\Data\\提取textbox内容.docx")
# 创建images数组
images = []
# 遍历文本框
for i in range(document.TextBoxes.Count):
textbox = document.TextBoxes.get_Item(i)
# 将textbox放入队列
nodes = queue.Queue()
nodes.put(textbox)
# 遍历nodes并获取node的子对象
while nodes.qsize() > 0:
node = nodes.get()
for i in range(node.ChildObjects.Count):
child = node.ChildObjects.get_Item(i)
# 判断node子对象是否为图片类型
# 是图片类型即加入images数组
if child.DocumentObjectType == DocumentObjectType.Picture:
picture = child if isinstance(child, DocPicture) else None
# 获取图片的字节并存储到images中
dataBytes = picture.ImageBytes
images.append(dataBytes)
# node子对象若不是图片类型且为复合对象则压入队列继续遍历
elif isinstance(child, ICompositeObject):
nodes.put(
child if isinstance(child, ICompositeObject) else None)
# 遍历images并保存图片
for i, item in enumerate(images):
fileName = "Image-{}.png".format(i)
with open("D:\\schedule\\Data\\" + fileName, 'wb') as imageFile:
imageFile.write(item)
document.Close()

使用 Spire.Doc for Python 获取文本框内表格内容的详细步骤如下:
from spire.doc import *
from spire.doc.common import *
# 创建Document对象
document = Document()
# 加载文件
document.LoadFromFile("D:\\schedule\\Data\\提取textbox内容.docx")
# 获取文档内第1个文本框
textbox = document.TextBoxes.get_Item(0)
# 计数图片
pic_count = 0
# 存放表格文本内容
tempStr = ''
# 遍历文本框内的对象
for j in range(textbox.ChildObjects.Count):
tb_obj = textbox.ChildObjects.get_Item(j)
# 若文本框的子对象为Table,则获取Table对象并转换为Table类型
if tb_obj.DocumentObjectType == DocumentObjectType.Table:
table = tb_obj if isinstance(tb_obj, Table) else None
# 遍历表格行及每行单元格
for m in range(table.Rows.Count):
row = table.Rows.get_Item(m)
for n in range(row.Cells.Count):
cell = row.Cells.get_Item(n)
# 遍历单元格中的段落
for k in range(cell.Paragraphs.Count):
paragraph = cell.Paragraphs.get_Item(k)
# 将段落文本保存到tempStr
tempStr += paragraph.Text + "\t"
# 遍历段落的子对象,若有图片,则在tempStr中定义图片名称
for p in range(paragraph.ChildObjects.Count):
para_obj = paragraph.ChildObjects.get_Item(p)
if para_obj.DocumentObjectType == DocumentObjectType.Picture:
tempStr += "Image-{}.png".format(pic_count) + "\t"
pic_count += 1
tempStr += "\r\n"
# 指定输出文件名并创建文件对象写入tempStr内容
outputFile = 'getTableFrTb.txt'
with open(outputFile, 'w') as fp:
fp.write(tempStr)
document.Close()

如果您希望删除结果文档中的评估消息,或者摆脱功能限制,请该Email地址已收到反垃圾邮件插件保护。要显示它您需要在浏览器中启用JavaScript。获取有效期 30 天的临时许可证。
Spire.Office for Java 9.4.0 已发布。在该版本中,Spire.Doc for Java 支持加载、操作和转换 Markdown 文档;Spire.PDF for Java 支持获取关键字的字体名和字体大小;Spire.XLS for Java 新增了一个转换工作表到 SVG 文档的方法;Spire.Barcode for Java 支持在二维码中间添加图片。此外,许多已知问题也在该版本中成功修复。详情请阅读以下内容。
获取 Spire.Office for Java 9.4.0 请点击:https://www.e-iceblue.cn/Downloads/Spire-Office-JAVA.html
新功能:
Document doc = new Document();
//load .md file
doc.loadFromFile("input.md");
//save to .md file
doc.saveToFile("output.md", com.spire.doc.FileFormat.Markdown);
//save to .docx file
//doc.saveToFile("output.docx", com.spire.doc.FileFormat.Docx);
//save to .doc file
//doc.saveToFile("output.doc", com.spire.doc.FileFormat.Doc);
//save to .pdf file
//doc.saveToFile("output.pdf", com.spire.doc.FileFormat.PDF);
doc.close();
Document doc = new Document();
//load .docx file
doc.loadFromFile("input.docx");
//load .doc file
//doc.loadFromFile("input.doc");
//save to .md file
doc.saveToFile("output.md", com.spire.doc.FileFormat.Markdown);
doc.close();问题修复:
新功能:
PdfDocument pdf = new PdfDocument();
pdf.loadFromFile(inputFile);
PdfPageBase page = pdf.getPages().get(0);
PdfTextFinder finds = new PdfTextFinder(page);
finds.getOptions().setTextFindParameter(EnumSet.of(TextFindParameter.IgnoreCase));
List<PdfTextFragment> result = finds.findAllText(page);
StringBuilder str = new StringBuilder();
for (PdfTextFragment find : result)
{
str.append("FontName:"+find.getTextStates()[0].getFontName());
str.append("FontSize:"+find.getTextStates()[0].getFontSize());
str.append("FontFamily:"+find.getTextStates()[0].getFontFamily());
str.append("Bold:"+find.getTextStates()[0].isBold());
str.append("Italic:"+find.getTextStates()[0].isItalic());
str.append("ForegroundColor:"+find.getTextStates()[0].getForegroundColor());
}PdfDocument doc = new PdfDocument();
doc.loadFromFile("input.pdf");
PdfTextReplaceOptions textReplaceOptions = new PdfTextReplaceOptions();
textReplaceOptions.setReplaceType(EnumSet.of(ReplaceActionType.Regex));
PdfPageBase page = doc.getPages().get(0);
PdfTextReplacer textReplacer = new PdfTextReplacer(page);
textReplacer.setOptions(textReplaceOptions);
String regularExpression = "\\bS\\w*L\\b";
textReplacer.replaceAllText(regularExpression, "NEW");
doc.saveToFile("output.pdf");
doc.dispose(); 问题修复:
新功能:
Workbook workbook = new Workbook();
workbook.loadFromFile("1.xlsx");
Worksheet sheet =workbook.getWorksheets().get(0);
FileOutputStream stream = new FileOutputStream(outputFile);
Dimension dimension=sheet.toSVGStream(stream, sheet.getFirstRow(), sheet.getFirstColumn(), sheet.getLastRow(), sheet.getLastColumn());
double Height=dimension.getHeight();
double Width=dimension.getWidth();问题修复:
问题修复:
新功能:
BarcodeSettings barCodeSetting = new BarcodeSettings();
BufferedImage image = ImageIO.read(new File("Image/1.png"));
barCodeSetting.setQRCodeLogoImage(image);问题修复:
Spire.XLS for Java 14.4.4 现已发布。该版本新增了一个 Excel 工作表转 SVG 文档的方法,支持返回 SVG 的尺寸。同时还成功修复了转换 Excel 文档到 PDF 文档和 HTML 文档时出现的问题。详情查看下文。
新功能:
Workbook workbook = new Workbook();
workbook.loadFromFile("1.xlsx");
Worksheet sheet =workbook.getWorksheets().get(0);
FileOutputStream stream = new FileOutputStream(outputFile);
Dimension dimension=sheet.toSVGStream(stream, sheet.getFirstRow(), sheet.getFirstColumn(), sheet.getLastRow(), sheet.getLastColumn());
double Height=dimension.getHeight();
double Width=dimension.getWidth();问题修复: