
以编程方式从 URL 下载 PDF 文件,对于构建 文档处理系统 、 网页爬虫 、内容聚合器或自动化报表生成器的开发者来说至关重要。自动化 PDF 下载与处理能够提升工作流效率,使开发者无需手动干预即可提取信息、归档文档或执行分析任务。
在本指南中,我们将演示如何使用 Python 配合 Spire.PDF 从 URL 下载 PDF,并实现纯内存处理、网络错误处理、大文件管理以及常见问题排查。
目录:
Spire.PDF for Python 支持 直接从内存加载 PDF ,无需依赖磁盘路径。这种内存处理方式速度更快,同时避免了不必要的磁盘 I/O 操作。
核心能力包括:
这些能力特别适用于 网页爬取流程 、 文档归档系统 、自动化报表生成 以及 内容提取工作流 ,因为这些场景通常对性能和内存效率要求较高。
使用 pip 安装 Spire.PDF 和 requests:
pip install spire.pdf requests
导入所需模块:
from spire.pdf import *
import requests
下面是一个完整示例,演示如何从 URL 下载 PDF、在内存中处理,并将其保存到本地磁盘。代码中的每一行都附带说明,便于理解。
import requests
from spire.pdf import *
def download_pdf_from_url():
# 指定 PDF URL
url = "https://example.com/sample.pdf"
# 发送 HTTP GET 请求下载 PDF
response = requests.get(url)
# 如果请求失败(4xx 或 5xx),则抛出异常
response.raise_for_status()
# 从下载的字节数据创建 Stream 对象
stream = Stream(response.content)
# 从 Stream 加载 PDF
document = PdfDocument(stream)
# 保存 PDF 到本地文件
document.SaveToFile("Downloaded.pdf")
document.Close()
print("PDF 下载并保存成功!")
if __name__ == "__main__":
download_pdf_from_url()
输出结果:

关键组件说明:
requests.get(url) —— 发送 HTTP GET 请求。服务器会返回响应头和 PDF 二进制数据。response.raise_for_status() —— 检查 HTTP 错误(例如 404、500)。response.content —— 包含原始 PDF 字节数据。Stream(response.content) —— 将字节数据包装为可读、可定位的内存流。PdfDocument(stream) —— 将 PDF 加载到内存中以进行后续操作。document.SaveToFile() —— 将 PDF 写入磁盘。该工作流会先将 PDF 数据加载到内存中再立即保存,从而提升处理速度并避免不必要的磁盘写入。
你可以直接在内存中提取元数据或文本,而无需将文件写入磁盘:
from spire.pdf import PdfTextExtractor
def process_pdf_from_url():
url = "https://example.com/sample.pdf"
response = requests.get(url)
response.raise_for_status()
# 在内存中加载 PDF
document = PdfDocument(Stream(response.content))
# 获取文档信息
print(f"页数: {document.Pages.Count}")
info = document.DocumentInformation
print(f"标题: {info.Title}")
print(f"作者: {info.Author}")
# 提取第一页文本
extractor = PdfTextExtractor(document.Pages[0])
text = extractor.ExtractText()
print(f"前100个字符: {text[:100]}")
document.Close()
if __name__ == "__main__":
process_pdf_from_url()
为什么这样有用: 你可以在不生成额外磁盘文件的情况下分析内容、建立文本索引或 提取元数据。这非常适合服务端脚本 、云函数 或批量处理任务 。
下载超大型 PDF(例如 100MB 以上)可能会占用大量内存。使用 流式下载 和临时文件可以降低内存消耗:
import tempfile
import os
def download_large_pdf(url: str, output_path: str):
try:
response = requests.get(url, stream=True, timeout=60)
response.raise_for_status()
# 将数据块写入临时文件
with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
tmp.write(chunk)
temp_path = tmp.name
# 从临时文件加载 PDF
document = PdfDocument()
document.LoadFromFile(temp_path)
document.SaveToFile(output_path)
document.Close()
# 清理临时文件
os.unlink(temp_path)
print(f"大型 PDF 已保存至: {output_path}")
except Exception as e:
print(f"错误: {e}")
说明:
stream=True 可避免一次性将整个文件加载到内存中。网络请求可能会间歇性失败。添加重试机制能够提升程序健壮性:
import time
def download_with_retry(url: str, output_path: str, max_retries: int = 3):
for attempt in range(max_retries):
try:
response = requests.get(url, timeout=30)
response.raise_for_status()
document = PdfDocument(Stream(response.content))
document.SaveToFile(output_path)
document.Close()
print(f"下载成功: {output_path}")
return True
except requests.exceptions.RequestException as e:
print(f"第 {attempt + 1} 次尝试失败: {e}")
if attempt < max_retries - 1:
wait_time = 2 ** attempt
print(f"{wait_time} 秒后重试...")
time.sleep(wait_time)
print("所有重试均失败。")
return False
为什么使用它: 指数退避(Exponential Backoff)可以避免对服务器造成过大压力,同时更优雅地处理临时网络故障。
问题: URL 未指向有效 PDF,导致返回 404 错误。
解决方案: 检查 URL 是否正确,并在必要时添加 User-Agent 请求头:
import requests
url = "https://example.com/missing.pdf"
headers = {'User-Agent': 'Mozilla/5.0'}
response = requests.get(url, headers=headers)
if response.status_code == 404:
print("PDF 未找到(404)")
问题: URL 返回的是 HTML 页面,而不是 PDF 文件。
解决方案: 检查 Content-Type,并解析 HTML 查找真实的 PDF 链接:
import requests
from bs4 import BeautifulSoup
url = "https://example.com/download-page"
response = requests.get(url)
content_type = response.headers.get('Content-Type', '')
if 'application/pdf' not in content_type and 'text/html' in content_type:
soup = BeautifulSoup(response.text, 'html.parser')
for link in soup.find_all('a', href=True):
if link['href'].endswith('.pdf'):
print(f"找到 PDF 链接: {link['href']}")
# 下载真实 PDF URL
问题: 文本提取结果不可读,通常是由于编码问题或扫描版 PDF 导致。
解决方案: 确保正确处理编码,或对扫描版 PDF 使用 OCR:
from spire.pdf import PdfDocument, PdfTextExtractor
document = PdfDocument("example.pdf")
extractor = PdfTextExtractor(document.Pages[0])
text = extractor.ExtractText()
print(text[:200])
# 如果文本仍然乱码,则 PDF 可能是图片型 PDF,可考虑使用 OCR
问题: 即使文件存在,document.Pages.Count 仍返回 0。
解决方案: PDF 可能已损坏或受到密码保护:
from spire.pdf import PdfDocument, Stream
with open("protected.pdf", "rb") as f:
pdf_bytes = f.read()
# 加载受密码保护的 PDF
document = PdfDocument(Stream(pdf_bytes), "password")
print(f"页数: {document.Pages.Count}")
本文演示了如何使用 Python 中的 Spire.PDF for Python 从 URL 下载 PDF 文件。借助 Stream 类,开发者可以直接从内存加载 PDF 数据,而无需不必要的磁盘 I/O,从而构建高效的文档处理流程。
我们介绍了完整工作流:使用 requests 库下载 PDF 数据、从字节数据创建 Stream 对象、加载 PdfDocument 实例、处理网络错误、管理大型文件以及排查常见问题。这些可直接用于生产环境的代码示例,为构建稳定可靠的 PDF 下载与处理系统提供了坚实基础。
若想完整体验 Spire.PDF for Python 的全部功能且不受评估版限制,你可以申请 30 天免费试用许可证。
Q1. 如何使用 Python 从 URL 下载 PDF?
使用 requests 库获取 PDF 数据,并通过 Spire.PDF 从内存加载:
response = requests.get(url)
stream = Stream(response.content)
document = PdfDocument(stream)
Q2. 如何处理需要身份验证的 PDF?
对于基础身份验证,可以使用 auth 参数:
response = requests.get(url, auth=('username', 'password'))
对于基于 Token 的身份验证,可添加请求头:
headers = {'Authorization': 'Bearer YOUR_TOKEN'}
response = requests.get(url, headers=headers)
Q3. 可以下载的 PDF 最大文件大小是多少?
理论限制取决于系统可用内存。对于大于 200MB 的文件,建议使用流式下载配合临时文件,而不是一次性全部加载到内存中。
Q4. 可以并行下载多个 PDF 吗?
可以。你可以使用 concurrent.futures 或 asyncio 同时下载多个 PDF,以提升性能。
from concurrent.futures import ThreadPoolExecutor
urls = ["url1.pdf", "url2.pdf", "url3.pdf"]
with ThreadPoolExecutor(max_workers=5) as executor:
executor.map(download_pdf, urls)
Spire.Presentation for Java 11.5.1 已发布。该版本新增了一个压缩图片的功能,并同时修复了一个在转换 PPT 到 PDF 时出现的问题。详情请查看以下内容。
新功能:
Presentation presentation = new Presentation();
presentation.loadFromFile(inputFile);
SlideCollection slides = presentation.getSlides();
for (int i = 0; i < slides.getCount(); i++) {
ISlide slide = slides.get(i);
ShapeCollection shapes = slide.getShapes();
for (int j = 0; j < shapes.getCount(); j++) {
IShape shape = shapes.get(j);
if (shape instanceof SlidePicture) {
SlidePicture slidepicture = (SlidePicture) shape;
// 压缩图片,目标分辨率50 DPI(数值越小压缩越大)
boolean result = slidepicture.getPictureFill().getCompressImage( true, 50f);
}
}
}
问题修复:
https://www.e-iceblue.cn/Downloads/Spire-Presentation-JAVA.html
Spire.Doc for Java 14.5.3 现已发布。该版本新增支持获取脚注或尾注的编号,同时支持“仅嵌入文档使用的字符”设置。此外,修复了多个 Word 转 PDF 相关问题,包括转换效果不一致以及图片模糊的问题。详情如下。
新功能:
Document doc = new Document();
doc.loadFromFile(inputFile);
StringBuilder sb = new StringBuilder();
for (int n = 0; n < doc.getSections().getCount(); n++) {
Section s = doc.getSections().get(n);
for (int i = 0; i < s.getParagraphs().getCount(); i++) {
Paragraph para = s.getParagraphs().get(i);
for (int j = 0, cnt = para.getChildObjects().getCount(); j < cnt; j++) {
ParagraphBase pBase = (ParagraphBase) para.getChildObjects().get(j);
if (pBase instanceof Footnote) {
Footnote fn = (Footnote) pBase;
if (fn.getFootnoteType() == FootnoteType.Footnote) {
StringBuilder fnText = new StringBuilder();
for (int k = 0; k < fn.getTextBody().getParagraphs().getCount(); k++) {
fnText.append(fn.getTextBody().getParagraphs().get(k).getText());
}
sb.append("Footnote:"+ fnText.toString() + "\nFootnoteID:" + fn.getId() + "\n");
}
if (fn.getFootnoteType() == FootnoteType.Endnote) {
StringBuilder enText = new StringBuilder();
for (int k = 0; k < fn.getTextBody().getParagraphs().getCount(); k++) {
enText.append(fn.getTextBody().getParagraphs().get(k).getText());
}
sb.append("Endnote:"+ enText.toString() + "\nEndnoteID:" + fn.getId() + "\n");
}
}
}
}
}
doc.setEmbedFontsInFile(true);
doc.setSaveSubsetFonts(true);
问题修复:
Spire.PDF 12.5.8 现已正式发布。该版本优化了从 PDF 到 Image 的转换功能。同时,一些在提取 PDF 页面文本时出现的问题也得以成功修复。更多详情如下。
问题修复:
Spire.Office for Python 11.5.0 已正式发布。该版本带来了多项强大功能提升:Spire.Doc 现在支持将 Word 文档转换为 Excel,Spire.PDF 在替换文本时可以指定文字颜色,Spire.XLS 实现了 Markdown 与 Excel 之间的无缝转换,Spire.OCR 则可与 AI 模型集成,提高识别准确率并增强图像文字识别能力。此外,所有核心组件(Spire.OCR除外)现在均支持 macOS ARM 通用架构。
除了这些新功能外,本版本还修复了大量与 Word、Excel、PDF 和 PowerPoint 文件的转换、处理及保存相关的已知问题,提供了更加稳定可靠的使用体验。更多详细信息请见下文。
https://www.e-iceblue.cn/Downloads/Spire-Office-Python.html
优化:
旧命名(已废弃):from spire.doc.charts import ChartType
新命名(请使用):from spire.doc.charts.ChartType import ChartType
新功能:
doc = Document()
doc.LoadFromFile(inputFile)
firstColumn = doc.Bookmarks["t_insert"].FirstColumn
lastColumn = doc.Bookmarks["t_insert"].LastColumn
doc = Document()
tableStyle = doc.Styles.Add(StyleType.TableStyle, "TestTableStyle3")
tableStyle.LeftIndent = 55
tableStyle.Borders.Color = Color.get_Green()
tableStyle.HorizontalAlignment = RowAlignment.Right
tableStyle.Borders.BorderType = BorderStyle.Single
section = doc.AddSection()
table = section.AddTable()
table.ResetCells(3, 3)
table.Rows[0].Cells[0].AddParagraph().AppendText("Aligned according to left indent")
table.PreferredWidth = PreferredWidth.FromPoints(300)
table.Format.StyleName = "TestTableStyle3"
style = doc.Styles.FindByName("TestTableStyle3")
if (style is not None) and isinstance(style, TableStyle):
tableStyle = style
tableStyle.Borders.Color = Color.get_Black()
tableStyle.Borders.BorderType = BorderStyle.Double
tableStyle.RowStripe = 3
tableStyle.ConditionalStyles[TableConditionalStyleType.OddRowStripe].Shading.BackgroundPatternColor = Color.get_LightBlue()
tableStyle.ConditionalStyles[TableConditionalStyleType.EvenRowStripe].Shading.BackgroundPatternColor = Color.get_LightCyan()
tableStyle.ColumnStripe = 1
tableStyle.ConditionalStyles[TableConditionalStyleType.EvenColumnStripe].Shading.BackgroundPatternColor = Color.get_LightPink()
table.ApplyStyle(tableStyle)
table.Format.StyleOptions = table.Format.StyleOptions | TableStyleOptions.ColumnStripe
doc.SaveToFile(outputFile, FileFormat.Docx)
style = doc.Styles.FindByName("TestTableStyle3")
style.RemoveSelf()
document = Document()
document.LoadFromFile("input.mhtml")
document.SaveToFile("output.pdf", FileFormat.PDF)
document.Close()
HtmlExportOptions options = doc.HtmlExportOptions
options.OfficeMathOutputMode = HtmlOfficeMathOutputMode.MathML
document = Document()
document.LoadFromFile("input.docx")
document.SaveToFile("output.xlsx", FileFormat.XLSX)
document.Close()
问题修复:
新功能:
workbook.HidePivotFieldList = true;
workbook = Workbook()
workbook.LoadFromFile(inputFile)
markdownOptions = MarkdownOptions()
markdownOptions.SavePicInRelativePath = False
markdownOptions.SaveHyperlinkAsRef = False
workbook.SaveToMarkdown(outputFile, markdownOptions)
workbook.Dispose()
workbook = Workbook()
workbook.LoadFromFile(inputFile)
wb.SaveToFile("out.xlsx", ExcelVersion.Version2010);
问题修复:
新功能:
问题修复:
新功能:
class MyPdfCustomAppearance(IPdfSignatureAppearance):
def __init__(self):
pass
def Generate(self, g: PdfCanvas):
x = 0.0
y = 0.0
fontSize = 10.0
font = PdfTrueTypeFont("SimSun", fontSize, PdfFontStyle.Regular, True)
lineHeight = fontSize
image = PdfImage.FromFile(inputImage)
g.DrawImage(image, x, y)
x = float(image.Width)
g.DrawString("Signer: Gary", font, PdfBrushes.get_Red(), PointF(x, y))
y += lineHeight + 5
g.DrawString("Phone: +86 12345678", font, PdfBrushes.get_Black(), PointF(x, y))
y += lineHeight + 5
g.DrawString("Address: Sichuan Province, China", font, PdfBrushes.get_Black(), PointF(x, y))
doc = PdfDocument()
doc.LoadFromFile(inputFile)
signatureMaker = PdfOrdinarySignatureMaker(doc, inputFile_pfx, "e-iceblue")
my_appearance = MyPdfCustomAppearance()
customAppearance = PdfCustomAppearance(my_appearance)
signatureMaker.MakeSignature("Signer", doc.Pages.get_Item(0), 90.0, 550.0, 270.0, 640.0, customAppearance)
doc.SaveToFile(outputFile)
doc.Close()
PdfDocument doc = new PdfDocument();
doc.LoadFromFile(inputFile);
// 定义一个矩形
RectangleF rctg = new RectangleF(0, 0, 200, 300);
var page = doc.Pages[0];
PdfTextFinder finder = new PdfTextFinder(page);
finder.Options.Parameter = TextFindParameter.None;
finder.Options.Area = rctg;
// 再矩形中查找特定文本
List<PdfTextFragment> findouts = finder.FindAllText();
StringBuilder sb = new StringBuilder();
foreach (PdfTextFragment find in findouts)
{
sb.AppendLine(find.Text);
sb.AppendLine(find.TextStates[0].FontName);
sb.AppendLine(find.TextStates[0].FontSize.ToString("F2"));
}
File.WriteAllText(outputFile, sb.ToString());
textReplacer.ReplaceText("文档", "文件", Color.get_Blue())
问题修复:
新功能:
新功能:
def _run_ai_test(self):
filename = "1.png"
output_file = "scan.txt"
file_path = r"F:\3.3.0AI\AI\ocr.xml"
model = "AIModel"
api_key = "ApiKey"
api_url = "ApiUrl"
self._update_ocr_config(file_path, model, api_key, api_url)
self._scan_img(filename, output_file)
def _scan_img(self, filename, output_file):
scanner = OcrScanner()
configure_options = ConfigureOptions()
configure_options.ModelPath = r"F:\3.3.0AI\AI"
configure_options.Language = "Japanese"
scanner.ConfigureDependencies(configure_options)
scanner.Scan(filename)
text = scanner.Text.ToString()
with open(output_file, "w", encoding="utf-8") as f:
f.write(text)
def _update_ocr_config(self, file_path, model, api_key, api_url):
tree = ET.parse(file_path)
root = tree.getroot()
model_node = root.find('./configs/model')
api_key_node = root.find('./configs/apiKey')
api_url_node = root.find('./configs/apiUrl')
if model_node is not None:
model_node.text = model
if api_key_node is not None:
api_key_node.text = api_key
if api_url_node is not None:
api_url_node.text = api_url
tree.write(file_path, encoding='utf-8', xml_declaration=True)
print("XML更新成功!")

以编程方式向 Word 文档插入数学公式对于构建科学文档生成器、学术报告系统、教育平台或工程自动化工具的开发者而言至关重要。无论是生成研究论文、技术文档还是数学工作表,自动化公式插入都能显著提高效率与一致性然而,在 Microsoft Word 中手动格式化公式耗时费力,而从头构建数学渲染引擎则极其复杂。开发者通常需要一种可靠的方式来在 Word 中添加公式,同时支持 LaTeX 和 MathML 等标准数学格式。
借助 Spire.Doc for Python,开发者可以通过直观的 API 直接从 LaTeX 和 MathML 代码向 Word 文档插入数学公式。本文演示了如何在 Python 中创建 Word 公式,包括插入公式、在 LaTeX、MathML 和 Office MathML(OMML)之间转换公式,以及将 Word 公式导出为不同的数学格式。
快速导航
Microsoft Word 使用 Office Math Markup Language(OMML) 作为其数学公式的内部格式。OMML 是一种基于 XML 的结构,用于控制 Word 文档中的公式布局、符号、分数、矩阵及其他数学元素。然而,直接创建或编辑 OMML 对大多数开发者而言较为繁琐。
在实际应用中,数学内容更常以 LaTeX 或 MathML 编写:
要以编程方式生成可编辑的 Word 公式,开发者通常需要在这些格式与 Word 原生公式对象之间进行转换。
Spire.Doc for Python 通过 OfficeMath 类提供对 Word 公式处理的本地支持。开发者无需手动生成 OMML 或依赖基于图像的变通方案,即可直接从 LaTeX 或 MathML 代码创建可编辑的 Word 公式。
主要功能包括:
| 功能 | 支持情况 |
|---|---|
| 从 LaTeX 插入公式 | ✓ |
| 从 MathML 插入公式 | ✓ |
| 将 Word 公式导出为 LaTeX | ✓ |
| 将 Word 公式导出为 MathML | ✓ |
| 访问原生 OMML 内容 | ✓ |
| 将公式渲染为图像 | ✓ |
这些功能对于学术报告生成、教育平台、MathML 转 Word 工作流、LaTeX 发布管道以及其他涉及数学内容的自动化文档生成场景尤为有用。
通过 pip 安装 Spire.Doc for Python:
pip install spire.doc
在 Python 脚本中导入所需的类:
from spire.doc import *
或者,也可以从 Spire.Doc for Python 下载页面手动安装该库。
LaTeX 是学术和科学文档中编写数学公式最常用的格式。借助 Spire.Doc for Python,可以将 LaTeX 表达式转换为原生 Word 公式对象,并直接将这些公式插入 DOCX 文件。
以下示例演示了如何使用 OfficeMath 类向 Word 文档插入多个 LaTeX 公式。
from spire.doc import *
def insert_latex_equations():
# 创建新的 Word 文档
doc = Document()
section = doc.AddSection()
# 添加标题段落
title_para = section.AddParagraph()
title_para.AppendText("从 LaTeX 插入的数学公式").CharacterFormat.FontName = "微软雅黑"
title_para.Format.HorizontalAlignment = HorizontalAlignment.Left
# 定义要插入的 LaTeX 公式
latex_equations = [
r"x = \frac{-b \pm \sqrt{b^2 - 4ac}}{2a}", # 二次方程公式
r"e^{i\pi} + 1 = 0", # 欧拉恒等式
r"\int_0^\infty e^{-x} \, dx = 1", # 定积分
# 求和公式
r"\sum_{i=1}^{n} i = \frac{n(n+1)}{2}",
r"\sum_{i=1}^{n} i = \frac{n(n+1)}{2}", # 求和公式
r"A = \begin{pmatrix} 1 & 2 \\ 3 & 4 \end{pmatrix}", # 矩阵
r"P(A \mid B) = \frac{P(B \mid A)P(A)}{P(B)}", # 概率公式
r"\sin^2\theta + \cos^2\theta = 1", # 三角恒等式
]
# 将每个 LaTeX 公式作为独立段落插入
for latex_code in latex_equations:
# 从 LaTeX 代码创建 OfficeMath 对象
office_math = OfficeMath(doc)
office_math.FromLatexMathCode(latex_code)
# 将公式添加到新段落
para = section.AddParagraph()
para.Items.Add(office_math)
# 保存文档
doc.SaveToFile("latex_equations.docx", FileFormat.Docx2019)
doc.Close()
print("LaTeX 公式插入成功!")
if __name__ == "__main__":
insert_latex_equations()
以下截图显示了生成的包含从 LaTeX 代码转换的公式的 Word 文档。

此方法支持复杂的 LaTeX 结构,如分数、积分、矩阵、希腊字母及其他数学运算符,同时保留原生 Word 公式格式。
除了独立公式外,还可以在文本段落中插入行内公式。这在句子或解释中嵌入数学表达式时非常有用。
from spire.doc import *
def insert_inline_equation():
# 创建新的 Word 文档
doc = Document()
section = doc.AddSection()
# 添加介绍性文本
para = section.AddParagraph()
para.AppendText("二次方程公式为 ").CharacterFormat.FontName = "微软雅黑"
# 插入行内公式
office_math = OfficeMath(doc)
office_math.FromLatexMathCode(r"x = \frac{-b \pm \sqrt{b^2 - 4ac}}{2a}")
para.Items.Add(office_math)
para.AppendText(",其中 a ≠ 0。").CharacterFormat.FontName = "微软雅黑"
# 保存文档
doc.SaveToFile("inline_equation.docx", FileFormat.Docx2019)
doc.Close()
if __name__ == "__main__":
insert_inline_equation()
插入的公式会以内联方式出现在文本中:

这种方法使得在常规文本内容中嵌入数学表达式变得简单,适用于教育材料、研究论文和技术文档。
如果需要将公式与格式化文本、标题、表格及其他结构化文档元素结合,还可以参考关于在 Python 中创建结构化 Word 文档的教程。
MathML(数学标记语言)是一种基于 XML 的标准,用于在网页和数字文档中表示数学表达式。它常用于在线教育平台、科学数据库和内容管理系统。以下示例展示了如何使用 Spire.Doc for Python 将 MathML 转换为 Word 公式。
from spire.doc import *
def insert_mathml_equations():
# 创建新的 Word 文档
doc = Document()
section = doc.AddSection()
# 添加标题段落
title_para = section.AddParagraph()
title_para.AppendText("来自 MathML 的数学公式").CharacterFormat.FontName = "微软雅黑"
# 定义要插入的 MathML 公式
mathml_equations = [
# 欧拉恒等式
r'<math xmlns="http://www.w3.org/1998/Math/MathML">'
r'<msup><mi>e</mi><mrow><mi>i</mi><mi>π</mi></mrow></msup>'
r'<mo>+</mo><mn>1</mn><mo>=</mo><mn>0</mn>'
r'</math>',
# 勾股定理
r'<math xmlns="http://www.w3.org/1998/Math/MathML">'
r'<msup><mi>a</mi><mn>2</mn></msup>'
r'<mo>+</mo>'
r'<msup><mi>b</mi><mn>2</mn></msup>'
r'<mo>=</mo>'
r'<msup><mi>c</mi><mn>2</mn></msup>'
r'</math>',
# 分数表达式
r'<math xmlns="http://www.w3.org/1998/Math/MathML">'
r'<mfrac>'
r'<mrow><mi>x</mi><mo>+</mo><mi>y</mi></mrow>'
r'<mrow><mi>z</mi><mo>−</mo><mn>1</mn></mrow>'
r'</mfrac>'
r'</math>',
# 积分方程
r'<math xmlns="http://www.w3.org/1998/Math/MathML">'
r'<msubsup><mo>∫</mo><mn>0</mn><mn>1</mn></msubsup>'
r'<msup><mi>x</mi><mn>2</mn></msup>'
r'<mi>d</mi><mi>x</mi>'
r'<mo>=</mo>'
r'<mfrac><mn>1</mn><mn>3</mn></mfrac>'
r'</math>'
]
# 将每个 MathML 公式作为独立段落插入
for mathml_code in mathml_equations:
# 从 MathML 代码创建 OfficeMath 对象
office_math = OfficeMath(doc)
office_math.FromMathMLCode(mathml_code)
# 将公式添加到新段落
para = section.AddParagraph()
para.Items.Add(office_math)
# 保存文档
doc.SaveToFile("mathml_equations.docx", FileFormat.Docx2019)
doc.Close()
print("MathML 公式插入成功!")
if __name__ == "__main__":
insert_mathml_equations()
以下截图显示了生成的包含从 MathML 代码转换的公式的 Word 文档。

MathML 支持在处理基于 XML 的教育内容、基于网页的公式系统以及以 MathML 格式存储数学表达式的 STEM 学习平台时特别有用。
可以在同一文档中混合使用 LaTeX 和 MathML 公式,从而在内容来源方面提供灵活性:
from spire.doc import *
def insert_mixed_equations():
# 创建新的 Word 文档
doc = Document()
section = doc.AddSection()
# 插入 LaTeX 公式
latex_para = section.AddParagraph()
latex_math = OfficeMath(doc)
latex_math.FromLatexMathCode(r"E = mc^2")
latex_para.Items.Add(latex_math)
# 插入 MathML 公式
mathml_para = section.AddParagraph()
mathml_math = OfficeMath(doc)
mathml_math.FromMathMLCode(
r'<math xmlns="http://www.w3.org/1998/Math/MathML">'
r'<mi>F</mi><mo>=</mo><mi>m</mi><mi>a</mi>'
r'</math>'
)
mathml_para.Items.Add(mathml_math)
# 保存文档
doc.SaveToFile("mixed_equations.docx", FileFormat.Docx2019)
doc.Close()
if __name__ == "__main__":
insert_mixed_equations()
当数学内容来自不同来源(如基于 LaTeX 的发布系统和基于 MathML 的 Web 应用程序)时,这种方法非常有用。
如果数学内容源自网页或基于 HTML 的系统,还可以参考关于在 Python 中将 HTML 内容转换为 Word 文档的教程。
除了向 Word 文档插入公式外,Spire.Doc for Python 还支持将 Word 公式导出为多种数学标记格式。这对于 Word、LaTeX 发布系统、基于 Web 的 MathML 平台以及自定义 XML 工作流之间的互操作性非常有用。
以下示例演示了如何从 Word 文档中提取公式并将其导出为 LaTeX、MathML 和 Office MathML(OMML)。
from spire.doc import *
def export_equation_formats():
# 加载包含公式的 Word 文档
doc = Document()
doc.LoadFromFile("equations.docx")
# 访问第一个段落
section = doc.Sections[0]
para = section.Paragraphs[0]
# 查找 OfficeMath 对象
for i in range(len(para.ChildObjects)):
item = para.ChildObjects[i]
if isinstance(item, OfficeMath):
# 导出为 LaTeX
latex_code = item.ToLaTexMathCode()
print("LaTeX:")
print(latex_code)
print()
# 导出为 MathML
mathml_code = item.ToMathMLCode()
print("MathML:")
print(mathml_code)
print()
# 导出为 Office MathML(OMML)
omml_code = item.ToOfficeMathMLCode()
print("OMML:")
print(omml_code)
# 将输出保存到文件
with open("equation.tex", "w", encoding="utf-8") as f:
f.write(latex_code)
with open("equation.xml", "w", encoding="utf-8") as f:
f.write(mathml_code)
with open("equation.omml", "w", encoding="utf-8") as f:
f.write(omml_code)
break
doc.Close()
if __name__ == "__main__":
export_equation_formats()
以下截图显示了在 Python 控制台中打印的导出的公式格式。

| 格式 | 主要用途 | 特点 |
|---|---|---|
| LaTeX | 学术出版和科学论文 | 紧凑的语法,在学术界广泛使用 |
| MathML | 基于 Web 的数学内容 | 基于 XML 的格式,专为浏览器和教育系统设计 |
| OMML | Microsoft Word 集成 | 原生 Office 公式格式,具有完整的 Word 兼容性 |
这些导出功能使得以下操作更加容易:
在某些场景中,可能需要将公式导出为图像文件,以便在演示文稿、网页或其他非可编辑上下文中使用。Spire.Doc for Python 允许将 Office Math 公式渲染为可以保存为图像文件的图像流。
from spire.doc import *
def render_equation_as_image():
# 创建包含公式的新 Word 文档
doc = Document()
section = doc.AddSection()
para = section.AddParagraph()
# 插入公式
office_math = OfficeMath(doc)
office_math.FromLatexMathCode(
r"\int_0^\infty e^{-x^2} dx = \frac{\sqrt{\pi}}{2}"
)
para.Items.Add(office_math)
# 将公式渲染为图像流
image_stream = office_math.SaveImageToStream(ImageType.Bitmap)
# 将图像保存到文件
with open("equations/equation.png", "wb") as f:
f.write(image_stream.ToArray())
# 释放非托管资源
image_stream.Dispose()
doc.Close()
print("公式渲染为图像成功!")
if __name__ == "__main__":
render_equation_as_image()
以下截图显示了渲染为图像文件的公式。

此功能特别适用于:
如果希望将整个 Word 文档渲染为图像而不是导出单个公式,请查看关于在 Python 中将 Word 文档转换为图像的教程。
在 Python 字符串中编写 LaTeX 代码时,始终使用原始字符串(前缀加 r)以防止转义序列被解释:
# 正确:使用原始字符串
latex_code = r"\int_0^\infty e^{-x} dx"
# 错误:反斜杠将被解释为转义序列
latex_code = "\int_0^\infty e^{-x} dx"
并非所有 LaTeX 命令都受 Word 公式引擎支持。某些高级 LaTeX 结构可能无法正确渲染。尽可能坚持使用标准数学 notation:
# 支持:标准数学 notation
office_math.FromLatexMathCode(r"\alpha + \beta = \gamma")
# 某些高级 LaTeX 结构可能不受支持
# office_math.FromLatexMathCode(r"\begin{align} ... \end{align}")
MathML 代码必须包含正确的命名空间声明才能正确解析:
# 正确:包含命名空间
mathml = r'<math xmlns="http://www.w3.org/1998/Math/MathML"><mi>x</mi></math>'
# 错误:缺少命名空间可能会失败
mathml = r'<math><mi>x</mi></math>'
处理完成后务必关闭文档以释放资源,尤其是在批量操作中:
doc = Document()
try:
# 处理公式
doc.SaveToFile("output.docx", FileFormat.Docx2019)
finally:
doc.Close() # 确保即使发生错误也能清理
将导出的 LaTeX 或 MathML 保存到文件时,确保特殊字符使用正确的 UTF-8 编码:
with open("equation.tex", "w", encoding="utf-8") as f:
f.write(latex_code)
使用后始终处置图像流以正确释放资源:
image_stream = office_math.SaveImageToStream(ImageType.Bitmap)
try:
with open("equation.png", "wb") as f:
f.write(image_stream.ToArray())
finally:
image_stream.Dispose()
本文演示了如何使用 Spire.Doc for Python 在 Python 中向 Word 文档插入数学公式。通过利用 Spire API,开发者可以从 LaTeX 和 MathML 代码创建 Word 公式,在 LaTeX、MathML 和 Word 原生 OMML 格式之间转换公式,并将公式渲染为图像。此功能对于自动化科学文档生成、教育内容创建和数学发布工作流至关重要。
Spire.Doc for Python 提供全面的公式处理能力,不仅限于基本插入,包括在 LaTeX 和 MathML 与 Word 原生 OMML 格式之间进行转换,以及将 Word 公式导出回 LaTeX、MathML 和 OMML。该库简化了复杂的数学排版,同时保持与 Microsoft Word 原生公式引擎的兼容性。
如果想评估 Spire.Doc for Python 的全部功能,可以申请 30 天免费许可证。
使用 Spire.Doc for Python 的 OfficeMath 类。创建 OfficeMath 对象,调用 FromLatexMathCode() 或 FromMathMLCode() 并传入公式代码,然后使用 para.Items.Add(office_math) 将其添加到段落。最后,使用 doc.SaveToFile() 保存文档。
可以。Spire.Doc for Python 支持使用 FromLatexMathCode() 方法从 LaTeX 代码插入公式。标准数学 notation(如分数、积分、上标、下标和希腊字母)可以转换为 Word 兼容公式。
支持。可以使用 FromMathMLCode() 方法从 MathML 创建 Word 公式。确保 MathML 内容包含正确的命名空间声明:
<math xmlns="http://www.w3.org/1998/Math/MathML">
可以。Spire.Doc for Python 提供 ToLaTexMathCode() 和 ToMathMLCode() 等方法,将 Office Math 公式导出为 LaTeX 或 MathML 格式。这对于内容迁移、存储或与其他数学系统集成非常有用。
在 OfficeMath 对象上使用 SaveImageToStream() 方法将公式渲染为图像流。然后可以将流保存为图像文件,并在演示文稿、网页或预览系统中使用。
Spire.PDF for Java 12.5.1 现已正式发布。该版本新增了 PDF 文档转换过程中的进度回调支持,同时修复了多个问题,例如 SVG 转 PDF 时内容不一致的问题。更多详情如下。
新功能:
class CustomProgressNotifier implements IProgressNotifier {
StringBuilder str = new StringBuilder();
private String outputFile;
public CustomProgressNotifier(String outputFile) {
this.outputFile = outputFile;
}
@Override
public void notify(float progress) {
str.append("==============Progress: ").append(progress).append("%==============\n");
try {
Files.writeString(Paths.get(outputFile), str.toString());
} catch (Exception e) {
e.printStackTrace();
}
}
}
//using
PdfDocument pdf=new PdfDocument();
pdf.loadFromFile("test.pdf");
pdf.registerProgressNotifier(new CustomProgressNotifier("Progress.txt"));
pdf.saveToFile("out.docx", FileFormat.DOCX);
pdf.dispose();
问题修复:

HTML 解析是 Java 开发中的一项关键任务,它能够帮助开发者提取结构化数据、分析内容,并与网页信息进行交互。无论你是在构建网页爬虫、验证 HTML 内容,还是从网页中提取文本与属性,一个可靠的工具都能大幅简化开发流程。
本文将介绍如何使用 Spire.Doc for Java 在 Java 中解析 HTML。它不仅提供强大的 HTML 解析能力,还能无缝结合文档处理功能,帮助开发者以更低代码量完成复杂任务。
虽然 Java 中有很多 HTML 解析库(例如 Jsoup),但 Spire.Doc 在“文档处理集成”和“低代码工作流”方面具有明显优势,非常适合追求开发效率的开发者。
以下是它在 Java HTML 解析场景中的几个优势:
直观的对象模型: HTML 转换为可遍历的文档结构(例如:节(Section)、段落(Paragraph)、表格(Table)),无需手动解析原始 HTML 标签。
全面的数据提取:可以轻松提取文本、属性、表格行/单元格,甚至样式(例如标题),无需额外的依赖项。
低代码工作流:只需少量代码即可加载并处理 HTML 内容,大幅降低开发成本。
轻量级集成:支持通过 Maven 或 Gradle 快速集成到 Java 项目中,依赖简单。
在 Java 中读取 HTML 前,请确保开发环境满足以下要求:
Maven 配置:在项目的 pom.xml 中添加仓库与依赖:
<repositories>
<repository>
<id>com.e-iceblue</id>
<name>e-iceblue</name>
<url>https://repo.e-iceblue.cn/repository/maven-public/</url>
</repository>
</repositories>
<dependencies>
<dependency>
<groupId>e-iceblue</groupId>
<artifactId>spire.doc</artifactId>
<version>14.9.5</version>
</dependency>
</dependencies>
添加后,Maven 会自动下载所需库及依赖。
如果你更倾向于手动安装,也可以从官网下载 JAR 文件后添加到项目中。
默认情况下,Spire.Doc 会在输出结果中添加评估版水印。如果希望移除水印并解锁完整功能,可以申请 30 天的免费试用许可证。
Spire.Doc 会将 HTML 解析为结构化对象模型,其中段落、表格、字段等元素都能作为 Java 对象进行访问。
下面通过几个实用示例演示如何提取 HTML 中的关键内容。
提取纯文本(不含 HTML 标签与格式)是内容索引、数据分析等场景中的常见需求。
下面示例演示如何解析 HTML 字符串并提取所有段落文本。
Java 代码:从 HTML 字符串中提取文本
import com.spire.doc.*;
import com.spire.doc.documents.*;
public class ExtractTextFromHtml {
public static void main(String[] args) {
// 定义 HTML 内容
String htmlContent = "<html>" +
"<body>" +
"<h1>Spire.Doc for Java 功能介绍" +
"<p>Spire.Doc for Java 是一款专业的 Word 文档处理组件,无需安装 Microsoft Office。</p>" +
"<ul>" +
"<li>支持创建、读取、修改 Word 文档</li>" +
"<li>支持 HTML 与 Word 相互转换</li>" +
"<li>支持 PDF、EPUB、图片等格式转换</li>" +
"</ul>" +
"<p>更多详情请访问:www.e-iceblue.com</p>" +
"</body>" +
"</html>";
// 创建 Document 对象
Document doc = new Document();
// 将 HTML 字符串解析到文档中
doc.addSection().addParagraph().appendHTML(htmlContent);
// 提取所有段落文本
StringBuilder extractedText = new StringBuilder();
for (Section section : (Iterable<Section>) doc.getSections()) {
for (Paragraph paragraph : (Iterable<Paragraph>) section.getParagraphs()) {
extractedText.append(paragraph.getText()).append("\n");
}
}
// 输出结果
System.out.println("Extracted Text:\n" + extractedText);
}
}
输出结果

HTML 表格通常用于存储结构化数据(例如商品列表、报表等)。Spire.Doc 将 <table> 标签解析为 Table 对象,因此可以非常方便地提取行列数据。
Java 代码:提取 HTML 表格行与单元格
import com.spire.doc.*;
import com.spire.doc.documents.*;
public class ExtractTableFromHtml {
public static void main(String[] args) {
// 包含表格的 HTML 内容
String htmlWithTable = "<html>" +
"<body>" +
"<h2>Spire.Doc for Java 产品价格表</h2>" +
"<table border='1'>" +
"<tr><th>产品版本</th><th>授权类型</th><th>价格(美元)</th></tr>" +
"<tr><td>标准版</td><td>单开发者授权</td><td>$599</td>" +
"<tr><td>专业版</td><td>团队授权</td><td>$1,499</td></tr>" +
"<tr><td>企业版</td><td>无限授权</td><td>$2,999</td></tr>" +
"</table>" +
"</body>" +
"</html>";
// 解析 HTML
Document doc = new Document();
doc.addSection().addParagraph().appendHTML(htmlWithTable);
// 提取表格数据
for (Section section : (Iterable<Section>) doc.getSections()) {
for (Object obj : section.getBody().getChildObjects()) {
if (obj instanceof Table) {
Table table = (Table) obj;
System.out.println("Table Data:");
// 遍历行
for (TableRow row : (Iterable<TableRow>) table.getRows()) {
// 遍历单元格
for (TableCell cell : (Iterable<TableCell>) row.getCells()) {
// 提取单元格文本
for (Paragraph para : (Iterable<Paragraph>) cell.getParagraphs()) {
System.out.print(para.getText() + "\t");
}
}
System.out.println();
}
}
}
}
}
}
输出结果:

通过 appendHTML() 方法将 HTML 字符串解析为 Word 文档后,你还可以利用 Spire.Doc 的 API 提取文本或图片等内容。
除了 HTML 字符串外,Spire.Doc for Java 还支持解析本地 HTML 文件和 Web URL,使其适用于各种真实业务场景。
要解析本地 HTML 文件,只需通过 loadFromFile(String filename, FileFormat.Html) 方法加载它即可进行处理。
Java 代码:读取并解析本地 HTML 文件
import com.spire.doc.*;
import com.spire.doc.documents.*;
public class ParseHtmlFile {
public static void main(String[] args) {
// 创建 Document 对象
Document doc = new Document();
// 加载 HTML 文件
doc.loadFromFile("input.html", FileFormat.Html);
// 提取文本
StringBuilder text = new StringBuilder();
for (Section section : (Iterable<Section>) doc.getSections()) {
for (Paragraph para : (Iterable<Paragraph>) section.getParagraphs()) {
text.append(para.getText()).append("\n");
}
}
System.out.println("Text from HTML File:\n" + text);
}
}
该示例会从加载的 HTML 文件中提取文本内容。如果你需要同时获取段落样式(例如 "Heading1"、"Normal"),可以使用 Paragraph.getStyleName() 方法。
输出结果:

你可能还需要:在 Java 中将 HTML 转换为 Word
在真实爬虫或网页分析场景中,你通常需要从在线网页中获取 HTML 内容。Spire.Doc 可以结合 Java 内置的 HttpClient(JDK 11+)来抓取网页 HTML 并进行解析。
Java 代码:抓取并解析网页 URL
import com.spire.doc.*;
import com.spire.doc.documents.*;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
public class ParseHtmlFromUrl {
// 可复用 HttpClient
private static final HttpClient httpClient = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(10))
.build();
public static void main(String[] args) {
String url = "https://business.segmentfault.com/about";
try {
// 获取 HTML 内容
System.out.println("Fetching from: " + url);
String html = fetchHtml(url);
// 使用 Spire.Doc 解析 HTML
Document doc = new Document();
Section section = doc.addSection();
section.addParagraph().appendHTML(html);
System.out.println("--- Headings ---");
// 提取标题
for (Paragraph para : (Iterable<Paragraph>) section.getParagraphs()) {
if (para.getStyleName() != null &&
para.getStyleName().startsWith("Heading")) {
System.out.println(para.getText());
}
}
} catch (Exception e) {
System.err.println("Error: " + e.getMessage());
}
}
// 获取 HTML 内容
private static String fetchHtml(String url) throws Exception {
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create(url))
.header("User-Agent", "Mozilla/5.0")
.timeout(Duration.ofSeconds(10))
.GET()
.build();
HttpResponse<String> response =
httpClient.send(request, HttpResponse.BodyHandlers.ofString());
// 检查请求状态
if (response.statusCode() != 200) {
throw new Exception("HTTP error: " + response.statusCode());
}
return response.body();
}
}
关键步骤:
HTTP 请求:使用 HttpClient 抓取网页 HTML,并通过 User-Agent 模拟浏览器访问,避免被服务器拦截。
HTML 解析:创建 Document 对象后,添加一个 Section 和一个 Paragraph,然后 appendHTML() 加载获取到的 HTML 内容。
内容提取:通过检查段落样式是否以 "Heading" 开头来提取网页标题。
输出结果:

借助 Spire.Doc for Java 库可以简化在 Java 中解析 HTML 的过程。借助它,你可以用最少的代码从 HTML 字符串、本地文件或 URL 中提取文本、表格和数据,而无需手动处理复杂 HTML 标签或引入沉重依赖。
无论你是在构建网页爬虫、分析 Web 内容,还是将 HTML 转换为其他格式(例如 HTML 转 PDF),Spire.Doc 都能帮助你快速构建稳定、高效的 HTML 处理流程。
答:这取决于你的需求:
答:Spire.Doc for Java 提供了专门的方式来处理不规范 HTML,使用 loadFromFile 方法配合 XHTMLValidationType.None 参数。此配置会禁用严格的 XHTML 验证,从而更宽容地解析不规范 HTML。
// 加载并解析格式错误的 HTML 文件
// 参数:文件路径,文件格式(HTML),验证类型(None)
doc.loadFromFile("input.html", FileFormat.Html, XHTMLValidationType.None);
不过,如果 HTML 结构严重损坏,仍可能导致解析失败。
答:可以。Spire.Doc 支持修改解析后的内容(例如编辑段落文本、删除表格行或新增元素),然后重新保存为 HTML。
// 将 HTML 解析到 Document 对象后:
Section section = doc.getSections().get(0);
Paragraph firstPara = section.getParagraphs().get(0);
firstPara.setText("Updated heading!"); // 修改文本
// 保存回 HTML
doc.saveToFile("modified.html", FileFormat.Html);
答:不需要,除非你是直接从 URL 加载 HTML。Spire.Doc 可以在没有网络连接的情况下解析本地文件或字符串中的 HTML。只有在从网页 URL 抓取 HTML 时,才需要网络连接来获取网页内容,而解析过程本身仍然可以离线完成。
Spire.XLS 16.5.6 现已发布。该版本增强了 Excel 到 EMF 的转换功能,还修复了两个在读取和计算公式时出现的问题。详情如下。
问题修复:

现代开发团队经常需要与不使用代码编辑器的项目经理、客户、审计人员或教育工作者分享 JavaScript 或 JSX 源代码。然而,原始的 .js 和 .jsx 文件在 VS Code 或 WebStorm 等工具之外难以查看,而手动将代码复制到 Word 文档中经常会破坏缩进、格式和可读性。
通过将 Spire.Doc for Python 与 Pygments 结合使用,开发人员可以在 Python 中将 JavaScript 转换为 Word,并实现语法高亮和可自定义的文档格式。这种自动化方法适用于技术文档、合规归档、教育材料、代码审查和客户交付物。
本文将介绍如何使用 Spire.Doc for Python 在 Python 中将 JavaScript 和 JSX 文件转换为 Word 文档,包括基本转换、高级格式化技术、批量处理和 PDF 导出。
快速导航
转换过程使用 Pygments 生成带语法高亮的 HTML,然后使用 Spire.Doc 的 HTML 导入功能将此 HTML 导入到 Word 文档中:
.js 或 .jsx 文件读取源代码highlight() 函数生成带语法高亮的 HTMLAppendHTML() 将 HTML 导入到 Word 中这种方法通过 Pygments 的内置样式提供语法着色,同时 Spire.Doc 处理文档结构,包括页边距、页眉、页脚和多格式导出。它为自动化转换过程提供了简单灵活的 API。
在 Python 中将 JavaScript 文件转换为 Word 文档之前,需要安装 Spire.Doc for Python 和 Pygments:
pip install spire.doc
pip install pygments
验证包是否可用:
import spire.doc
from pygments import highlight
from pygments.formatters import HtmlFormatter
或者,可以下载 Spire.Doc for Python 并将其添加到项目中。
以下示例将 JavaScript 文件转换为带有语法高亮的 Word 文档:
from spire.doc import *
from pygments import highlight
from pygments.lexers import JavascriptLexer
from pygments.formatters import HtmlFormatter
def convert_js_to_word(input_file: str, output_file: str) -> None:
"""将 JavaScript 文件转换为带有语法高亮的 Word 文档。"""
with open(input_file, "r", encoding="utf-8") as file:
js_code = file.read()
document = Document()
section = document.AddSection()
section.PageSetup.Margins.All = 50
title_paragraph = section.AddParagraph()
title_text = title_paragraph.AppendText(f"源代码:{input_file}")
title_text.CharacterFormat.FontName = "微软雅黑"
title_text.CharacterFormat.FontSize = 14
title_text.CharacterFormat.Bold = True
title_paragraph.Format.AfterSpacing = 10
html_formatter = HtmlFormatter(
nowrap=True,
style='colorful',
noclasses=True
)
highlighted_html = highlight(js_code, JavascriptLexer(), html_formatter)
code_paragraph = section.AddParagraph()
code_paragraph.AppendHTML(f'<pre style="font-family: Consolas; font-size: 10pt;">{highlighted_html}</pre>')
document.SaveToFile(output_file, FileFormat.Docx)
document.Close()
print(f"已将 {input_file} 转换为 {output_file}")
convert_js_to_word("app.js", "JavaScriptCode.docx")

noclasses=True)Spire.Doc 可以导入由 Pygments 生成的带语法高亮的 HTML,使 JavaScript 代码格式和颜色能够在 Word 文档中保留。
对于 JSX 文件,建议使用 JsxLexer 而不是 JavascriptLexer,以实现对组件标签和嵌入式 JSX 表达式更准确的语法高亮。
JSX 输入示例(App.jsx):
import React, { useState } from 'react';
const TodoList = () => {
const [todos, setTodos] = useState([]);
return (
<div className="todo-container">
<h1>我的任务</h1>
</div>
);
};
export default TodoList;
在生成带语法高亮的 HTML 时使用 JsxLexer:
from pygments.lexers import JsxLexer
highlighted_html = highlight(
jsx_code,
JsxLexer(),
html_formatter
)
然后使用相同的 AppendHTML() 工作流将高亮的 JSX 内容转换为 Word:
convert_js_to_word("App.jsx", "ReactComponent.docx")
转换结果如下所示:

与标准 JavaScript 词法分析器相比,JsxLexer 对 JSX 标签、属性和嵌入式表达式的识别能力更强,从而在生成的 Word 文档中实现更准确的语法着色。
如果需要转换大量 JavaScript 或 JSX 文件,可以通过扫描文件夹并批量生成 Word 文档来自动化此过程。
import os
from pathlib import Path
def batch_convert_js_files(source_folder: str, output_folder: str) -> None:
"""将文件夹中的所有 JavaScript 文件转换为 Word 文档。"""
Path(output_folder).mkdir(parents=True, exist_ok=True)
js_extensions = ('.js', '.jsx', '.mjs')
converted_count = 0
error_count = 0
for filename in os.listdir(source_folder):
if filename.lower().endswith(js_extensions):
input_path = os.path.join(source_folder, filename)
base_name = os.path.splitext(filename)[0]
output_path = os.path.join(output_folder, f"{base_name}.docx")
try:
convert_js_to_word(input_path, output_path)
converted_count += 1
except Exception as e:
print(f"转换 {filename} 时出错:{str(e)}")
error_count += 1
print(f"\n批量转换完成:")
print(f" 已转换:{converted_count} 个文件")
print(f" 错误:{error_count} 个文件")
batch_convert_js_files("src/scripts", "output/docs")
行号可以提高代码审查、审计或技术文档期间的可读性。由于 Word HTML 渲染可能不完全支持 Pygments 的内置行号布局,一种实用的方法是在语法高亮后 prepending 自定义行号。
html_formatter = HtmlFormatter(
nowrap=True,
noclasses=True,
style="colorful"
)
highlighted_html = highlight(
js_code,
JavascriptLexer(),
html_formatter
)
highlighted_lines = highlighted_html.splitlines()
numbered_lines = []
for index, line in enumerate(highlighted_lines, start=1):
numbered_line = (
f'<span style="color: gray; font-weight: bold;">'
f'{index:4d} '
f'</span>{line}'
)
numbered_lines.append(numbered_line)
combined_html = (
'<pre style="font-family: Consolas; '
'font-size: 10pt; line-height: 1.4;">'
+ '\n'.join(numbered_lines) +
'</pre>'
)
paragraph.AppendHTML(combined_html)
生成的带行号的 Word 文档如下所示:

页眉和页脚通过添加标题、页码和文档元数据来帮助组织生成的 Word 文档。这对于正式报告或导出的技术文档特别有用。
def add_document_metadata(section: Section, document_title: str) -> None:
"""向文档节添加页眉和页脚。"""
header = section.HeadersFooters.Header.AddParagraph()
header_text = header.AppendText(document_title)
header_text.CharacterFormat.FontName = "微软雅黑"
header_text.CharacterFormat.FontSize = 10
header_text.CharacterFormat.TextColor = Color.get_Black()
header.Format.HorizontalAlignment = HorizontalAlignment.Left
header.Format.TextAlignment = TextAlignment.Top
header.Format.Borders.Bottom.BorderType = BorderStyle.Single
header.Format.Borders.Bottom.Color = Color.get_Black()
footer = section.HeadersFooters.Footer.AddParagraph()
footer.Format.HorizontalAlignment = HorizontalAlignment.Center
footer.Format.TextAlignment = TextAlignment.Bottom
page_field = footer.AppendField("page", FieldType.FieldPage)
page_field.CharacterFormat.FontName = "微软雅黑"
page_field.CharacterFormat.FontSize = 9
footer.AppendText(" / ")
total_pages_field = footer.AppendField("numPages", FieldType.FieldNumPages)
total_pages_field.CharacterFormat.FontName = "微软雅黑"
total_pages_field.CharacterFormat.FontSize = 9
document = Document()
document.LoadFromFile("CodeWithLines.docx")
section = document.Sections[0]
add_document_metadata(section, "JavaScript 源代码文档")
document.SaveToFile("CodeWithHeadersFooters.docx", FileFormat.Docx)
生成的带页眉和页脚的 Word 文档如下所示:

有关更多高级自定义选项,请参阅关于如何在 Python 中向 Word 文档添加页眉和页脚的指南。
除了 DOCX 输出外,Spire.Doc 还可以将带语法高亮的 JavaScript 代码直接导出为 PDF 格式。这在分发只读文档或在 Microsoft Word 环境之外共享代码时非常有用。
def convert_js_to_pdf(input_file: str, output_file: str) -> None:
"""将 JavaScript 文件直接转换为 PDF。"""
with open(input_file, "r", encoding="utf-8") as file:
js_code = file.read()
document = Document()
section = document.AddSection()
section.PageSetup.Margins.All = 50
html_formatter = HtmlFormatter(noclasses=True, style='colorful')
highlighted_html = highlight(js_code, JavascriptLexer(), html_formatter)
paragraph = section.AddParagraph()
paragraph.AppendHTML(f'<pre style="font-family: Consolas; font-size: 10pt;">{highlighted_html}</pre>')
document.SaveToFile(output_file, FileFormat.PDF)
document.Close()
convert_js_to_pdf("app.js", "JavaScriptCode.pdf")
有关更多高级 PDF 转换技术,包括布局控制和文档格式,请参阅关于在 Python 中将 Word 文档转换为 PDF的详细指南。
Pygments 提供多种内置配色方案:
def convert_with_custom_style(input_file: str, output_file: str, style_name: str = 'monokai') -> None:
"""使用自定义高亮样式将 JavaScript 转换为 Word。"""
with open(input_file, "r", encoding="utf-8") as file:
js_code = file.read()
document = Document()
section = document.AddSection()
section.PageSetup.Margins.All = 50
html_formatter = HtmlFormatter(
noclasses=True,
style=style_name,
nowrap=True
)
highlighted_html = highlight(js_code, JavascriptLexer(), html_formatter)
paragraph = section.AddParagraph()
paragraph.AppendHTML(f'<pre style="font-family: Consolas; font-size: 10pt;">{highlighted_html}</pre>')
document.SaveToFile(output_file, FileFormat.Docx)
document.Close()
convert_with_custom_style("app.js", "CodeMonokai.docx", style_name='monokai')
可用样式包括:'monokai'、'colorful'、'vim'、'vs'、'tango'、'friendly'、'default'
问题: 默认的 HtmlFormatter 生成 CSS 类而不是内联样式,如果没有外部样式表,Word 无法处理。
解决方案: 始终使用 noclasses=True:
html_formatter = HtmlFormatter(noclasses=True, style='colorful')
highlighted_html = highlight(js_code, JavascriptLexer(), html_formatter)
问题: 在不指定 UTF-8 编码的情况下读取文件会导致某些平台上的字符损坏。
解决方案: 显式指定 UTF-8 编码:
with open(input_file, "r", encoding="utf-8") as file:
js_code = file.read()
对于带有 BOM(字节顺序标记)的文件,使用 utf-8-sig:
with open(input_file, "r", encoding="utf-8-sig") as file:
js_code = file.read()
问题: 未将高亮代码包装在 <pre> 标签中会导致缩进消失。
解决方案: 将语法高亮的 HTML 包装在 <pre> 标签中:
highlighted_html = highlight(js_code, JavascriptLexer(), html_formatter)
paragraph.AppendHTML(f'<pre style="font-family: Consolas;">{highlighted_html}</pre>')
问题: 当前 Python 环境中未安装包。
解决方案:
pip install spire.doc
对于虚拟环境,在安装前确保激活:
source venv/bin/activate # Linux/Mac
venv\Scripts\activate # Windows
pip install spire.doc
问题: 非常大的 JavaScript 文件(10,000+ 行)可能导致转换缓慢。
解决方案: 分块处理文件:
def convert_large_file(input_file: str, output_file: str, chunk_size: int = 500) -> None:
"""分块转换大型 JavaScript 文件。"""
with open(input_file, "r", encoding="utf-8") as file:
lines = file.readlines()
document = Document()
section = document.AddSection()
section.PageSetup.Margins.All = 50
html_formatter = HtmlFormatter(noclasses=True, style='colorful')
for i in range(0, len(lines), chunk_size):
chunk = ''.join(lines[i:i + chunk_size])
highlighted_html = highlight(chunk, JavascriptLexer(), html_formatter)
paragraph = section.AddParagraph()
paragraph.AppendHTML(f'<pre style="font-family: Consolas; font-size: 10pt;">{highlighted_html}</pre>')
document.SaveToFile(output_file, FileFormat.Docx)
document.Close()
本文演示了如何使用 Spire.Doc for Python 和 Pygments 在 Python 中将 JavaScript 和 JSX 文件转换为 Word 文档。通过利用带有 HtmlFormatter 的 highlight() 函数和 Spire.Doc 的 AppendHTML() 方法,开发人员可以使用语法高亮自动化代码文档工作流。
Spire.Doc for Python 提供文档生成功能,包括表格创建、图像插入、页眉/页脚管理和多格式导出。
可以申请30 天免费许可证来评估所有功能。
可以。Pygments 可以使用 JavaScript 词法分析器高亮显示许多 JSX 构造,包括组件标签、props 和嵌入式表达式。但是,JSX 特定语法可能不会获得专用的高亮类别。
不需要。Spire.Doc for Python 独立运行,无需 Microsoft Word。该库直接生成 DOCX 文件,使其适用于服务器环境和 CI/CD 管道。
可以。Spire.Doc 支持多种导出格式:
document.SaveToFile("output.pdf", FileFormat.PDF)
document.SaveToFile("output.html", FileFormat.Html)
document.SaveToFile("output.rtf", FileFormat.Rtf)
使用 TypescriptLexer:
from pygments.lexers import TypescriptLexer
highlighted_html = highlight(ts_code, TypescriptLexer(), html_formatter)
适合。Python 自动化可以与 CI/CD 管道和批处理工作流集成。本地执行避免了将源代码上传到在线转换器的安全风险。对于大型部署,考虑实施日志记录、进度报告和错误跟踪。
可以。Pygments 提供众多内置样式:
html_formatter = HtmlFormatter(noclasses=True, style='monokai')
可用样式:'monokai'、'colorful'、'vim'、'vs'、'tango'、'friendly'、'default'