☰
PaddleOCR PP-Structure 文档恢复实战:开启 return_word_box 让识别结果返回每个文字的位置
2026/10/3 13:40:18 网站建设 项目流程

PaddleOCR PP-Structure 文档恢复实战:开启 return_word_box 让识别结果返回每个文字的位置

【免费下载链接】PaddleOCRTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.项目地址: https://gitcode.com/GitHub_Trending/pa/PaddleOCR

本指南围绕 PaddleOCR 仓库中ppstructure文档恢复(Document Recovery)管线,讲解如何通过--return_word_box=True让识别模型在返回文本内容的同时,输出每个单词/单字的精确位置框,并配套英文、中文两类文档的完整落地步骤。读完本文,你将掌握return_word_box的开关方式、输出 JSON 的字段语义、从识别解码到单词框计算的底层原理,以及如何将位置信息用于版面恢复、Markdown/Docx 生成与结构化数据抽取。

一、为什么文档恢复需要"文字位置"

在横向排版(横排)的文档中,PP-Structure 的识别模型不仅能输出识别出的文字内容,还能输出每一个文字/单词所在的坐标位置。这一能力对文档恢复(Document Recovery)与 AI 结构化数据处理至关重要:

  • 版面恢复:恢复出的 Word/Markdown 需要还原原文的阅读顺序、段落缩进与分栏结构,而这一切都依赖每个文字块的空间坐标;
  • 段落合并:PP-Structure 依据文字框的 x 坐标偏移判断段落首行缩进(见下文check_merge_method逻辑),没有位置信息就无法正确分段;
  • 结构化输出:将 OCR 结果交给 LLM 或下游系统时,带坐标的文本(所谓 "grounding")可以定位"这句话在文档的哪个区域",显著提升检索与引用价值。

return_word_box正是控制该能力的核心开关,默认关闭,开启后识别结果的每个文本行会额外携带text_word(单词/单字内容)与text_word_region(单词/单字级坐标框)两个字段。

二、return_word_box输出结构:一行文字返回了什么

从 predict_system.py 中_predict_text的实现可以看到,开启开关后每个文本行返回的 dict 从 3 个字段扩展为 5 个字段:

字段含义开启前开启后
text整行识别文本✔✔
confidence整行置信度✔✔
text_region整行的四点检测框([x1,y1],[x2,y1],[x2,y2],[x1,y2]形式)✔✔
text_word单词/单字内容列表,如["Hello", "world"]或中文字符序列✘✔
text_word_region与text_word一一对应的单词/单字级四点框列表✘✔

关键实现片段(ppstructure/predict_system.py):

if self.return_word_box: word_box_content_list, word_box_list = cal_ocr_word_box( rec_str, box, rec_res[2] ) res.append({ "text": rec_str, "confidence": float(rec_conf), "text_region": box.tolist(), "text_word": word_box_content_list, "text_word_region": word_box_list, }) else: res.append({ "text": rec_str, "confidence": float(rec_conf), "text_region": box.tolist(), })

也就是说,return_word_box并不改变识别模型本身,而是在后处理阶段基于整行检测框 + 识别解码时的逐字对齐信息,把整行框"拆"成单词/单字级别的子框。

三、源码原理:单词框是如何算出来的

3.1 识别解码阶段的逐字信息

单词框的源头在识别后处理。在 rec_postprocess.py 中,BaseRecLabelDecode.decode支持return_word_box参数:当开启时,会把解码得到的字符索引序列转成word_list(单词内容)、word_col_list(单词覆盖的列区间)、state_list(每个单词属于cn中文还是其他语言)三份信息,作为识别结果的第三个元素返回:

if return_word_box: word_list, word_col_list, state_list = self.get_word_info(text, selection) result_list.append(( text, np.mean(conf_list).tolist(), [len(text_index[batch_idx]), word_list, word_col_list, state_list], ))

对于 CTC 解码路径,CTCLabelDecode 还会用实际宽高比修正列数,保证列区间与真实图像坐标对齐。因此rec_res[2]中携带的是"解码层面"的逐字对齐信息,这是后续坐标换算的数据基础。

3.2 从列区间到像素坐标:cal_ocr_word_box

真正把"解码列区间"换算成"图像像素框"的是 ppstructure/utility.py 中的cal_ocr_word_box:

col_num, word_list, word_col_list, state_list = rec_word_info cell_width = (bbox_x_end - bbox_x_start) / col_num

其换算思路是:

  1. 用整行检测框的宽度除以解码总列数,得到每个字符列的像素宽度cell_width;
  2. 对于英文等非中文单词(state != "cn"):直接把单词覆盖的列区间映射回像素区间,得到矩形四点框;
  3. 对于中文单字(state == "cn"):先统计已切分出的汉字宽度取均值avg_char_width,再以每个字的列中心为中心、向两侧各延伸半个平均字宽,生成单字框——这是对连续中文缺乏显式切分时的近似策略。
center_x = (center_idx + 0.5) * cell_width cell_x_start = max(int(center_x - avg_char_width / 2), 0) + bbox_x_start cell_x_end = (min(int(center_x + avg_char_width / 2), bbox_x_end - bbox_x_start) + bbox_x_start)

3.3 完整调用链

一次开启return_word_box的文档恢复推理,单词框计算的完整链路为:

predict_system.py (StructureSystem.__call__) └─ _predict_text → TextSystem 检测+识别 └─ rec_postprocess.py decode(return_word_box=True) → 逐字列信息 └─ cal_ocr_word_box(rec_str, box, rec_res[2]) → 单词/单字像素框 └─ draw_structure_result(vis_font_path) → 绘制整行框 + 单词框可视化 └─ (recovery=True) recovery_to_doc / recovery_to_markdown → 还原 docx/markdown

其中draw_structure_result(ppstructure/utility.py)在可视化时会把text_word_region中的每个子框都画出来(跳过宽或高为 0 的异常框),这就是结果图中看到"整行大框 + 行内单词小框"双层框的来源。

四、实战一:英文文档恢复(返回单词位置)

4.1 下载四个推理模型

在仓库ppstructure/目录下执行以下命令,下载英文检测、识别、表格结构、版面分析四类推理模型:

cd PaddleOCR/ppstructure ## download model mkdir inference && cd inference ## Download the detection model of the ultra-lightweight English PP-OCRv3 model and unzip it wget https://paddleocr.bj.bcebos.com/PP-OCRv3/english/en_PP-OCRv3_det_infer.tar && tar xf en_PP-OCRv3_det_infer.tar ## Download the recognition model of the ultra-lightweight English PP-OCRv3 model and unzip it wget https://paddle-model-ecology.bj.bcebos.com/paddlex/official_inference_model/paddle3.0.0/en_PP-OCRv3_mobile_rec_infer.tar && tar xf en_PP-OCRv3_mobile_rec_infer.tar ## Download the ultra-lightweight English table inch model and unzip it wget https://paddleocr.bj.bcebos.com/ppstructure/models/slanet/paddle3.0b2/en_ppstructure_mobile_v2.0_SLANet_infer.tar tar xf en_ppstructure_mobile_v2.0_SLANet_infer.tar ## Download the layout model of publaynet dataset and unzip it wget https://paddleocr.bj.bcebos.com/ppstructure/models/layout/picodet_lcnet_x1_0_fgd_layout_infer.tar tar xf picodet_lcnet_x1_0_fgd_layout_infer.tar cd ..

4.2 执行推理

在ppstructure/目录下运行(注意将--image_dir替换为你自己的英文文档图片路径):

python predict_system.py \ --image_dir=./docs/ppstructure/images/table_1.png \ --det_model_dir=inference/en_PP-OCRv3_det_infer \ --rec_model_dir=inference/en_PP-OCRv3_mobile_rec_infer \ --rec_char_dict_path=../ppocr/utils/en_dict.txt \ --table_model_dir=inference/en_ppstructure_mobile_v2.0_SLANet_infer \ --table_char_dict_path=../ppocr/utils/dict/table_structure_dict.txt \ --layout_model_dir=inference/picodet_lcnet_x1_0_fgd_layout_infer \ --layout_dict_path=../ppocr/utils/dict/layout_dict/layout_publaynet_dict.txt \ --vis_font_path=../doc/fonts/simfang.ttf \ --recovery=True \ --output=../output/ \ --return_word_box=True

推理完成后,在../output/structure/table_1/show_0.jpg查看可视化结果:每个版面区域被不同颜色框标注,文字区域内部可以看到整行文字框 + 行内逐词小框的双层结构,这正是return_word_box生效的直观体现。

五、实战二:中文文档恢复(返回单字位置)

中文与英文的差异在于:英文按空格分词,每个单词一个框;中文没有空格,cal_ocr_word_box会按平均字宽为每个单字生成一个框(state == "cn"分支)。

5.1 下载四个推理模型

cd PaddleOCR/ppstructure ## download model cd inference ## Download the detection model of the ultra-lightweight Chinese PP-OCRv3 model and unzip it wget https://paddle-model-ecology.bj.bcebos.com/paddlex/official_inference_model/paddle3.0.0/PP-OCRv3_mobile_det_infer.tar && tar xf PP-OCRv3_mobile_det_infer.tar ## Download the recognition model of the ultra-lightweight Chinese PP-OCRv3 model and unzip it wget https://paddle-model-ecology.bj.bcebos.com/paddlex/official_inference_model/paddle3.0.0/PP-OCRv3_mobile_rec_infer.tar && tar xf PP-OCRv3_mobile_rec_infer.tar ## Download the ultra-lightweight Chinese table inch model and unzip it wget https://paddleocr.bj.bcebos.com/ppstructure/models/slanet/paddle3.0b2/ch_ppstructure_mobile_v2.0_SLANet_infer.tar tar xf ch_ppstructure_mobile_v2.0_SLANet_infer.tar ## Download the layout model of CDLA dataset and unzip it wget https://paddleocr.bj.bcebos.com/ppstructure/models/layout/picodet_lcnet_x1_0_fgd_layout_cdla_infer.tar tar xf picodet_lcnet_x1_0_fgd_layout_cdla_infer.tar cd ..

中文场景使用 CDLA 版面数据集训练的版面模型,layout_dict_path相应换成 CDLA 字典。

5.2 准备测试图片

下载/上传一张中文文档截图作为测试输入(即下方示例中的2.png,也可替换为任意中文论文、合同或报告截图):

5.3 执行推理

python predict_system.py \ --image_dir=./docs/table/2.png \ --det_model_dir=inference/PP-OCRv3_mobile_det_infer \ --rec_model_dir=inference/PP-OCRv3_mobile_rec_infer \ --rec_char_dict_path=../ppocr/utils/ppocr_keys_v1.txt \ --table_model_dir=inference/ch_ppstructure_mobile_v2.0_SLANet_infer \ --table_char_dict_path=../ppocr/utils/dict/table_structure_dict_ch.txt \ --layout_model_dir=inference/picodet_lcnet_x1_0_fgd_layout_cdla_infer \ --layout_dict_path=../ppocr/utils/dict/layout_dict/layout_cdla_dict.txt \ --vis_font_path=../doc/fonts/chinese_cht.ttf \ --recovery=True \ --output=../output/ \ --return_word_box=True

推理完成后在../output/structure/2/show_0.jpg查看可视化。中文文档的每个文本块内部可以看到逐字级的细框(红色、绿色、蓝色等不同颜色区分不同版面区域),覆盖了原文的段落、图注与公式区域:

六、参数详解与默认值对照

上述命令中的关键参数在 ppstructure/utility.py 的init_args中定义,整理如下:

参数类型默认值说明
--image_dirstr必填输入图片或图片目录,支持多图
--det_model_dirstr无文本检测推理模型目录
--rec_model_dirstr无文本识别推理模型目录
--rec_char_dict_pathstr无识别字典,英文用en_dict.txt,中文用ppocr_keys_v1.txt
--table_model_dirstr无表格结构推理模型目录(SLANet)
--table_char_dict_pathstr../ppocr/utils/dict/table_structure_dict_ch.txt表格结构字典
--layout_model_dirstr无版面分析推理模型目录
--layout_dict_pathstr../ppocr/utils/dict/layout_dict/layout_publaynet_dict.txt版面类别字典(中文用layout_cdla_dict.txt)
--layout_score_thresholdfloat0.5版面框置信度阈值
--layout_nms_thresholdfloat0.5版面框 NMS 阈值
--vis_font_pathstr无可视化绘制字体,英文推荐simfang.ttf,中文推荐chinese_cht.ttf
--recoveryboolFalse是否启用版面恢复(还原 docx/markdown)
--recovery_to_markdownboolFalse是否额外输出 Markdown 文件
--outputstr./output结果输出根目录
--return_word_boxboolFalse是否返回每个单词/单字的位置框
--modestrstructure推理模式,可选structure/kie
--layoutboolTrue是否启用版面分析
--tableboolTrue表格区域是否走表格识别
--ocrboolTrue非表格区域是否走 OCR
--image_orientationboolFalse是否启用图像方向识别(自动旋转 90/180/270 度)

注意:--layout=False时--ocr会被强制置为 False(见 predict_system.py 的告警逻辑),二者存在联动约束。

七、结果文件与后续处理

开启--recovery=True后,每个图片会在output/structure/<图片名>/目录下生成:

  • show_0.jpg:双层框可视化结果图(整行框 + 单词/单字框);
  • res_0.txt:结构化 JSON 结果,每个版面区域一行,文本区域包含text、confidence、text_region、text_word、text_word_region字段(由save_structure_res写出,见 predict_system.py);
  • <图片名>_ocr.docx:恢复出的 Word 文档(recovery_to_doc.py);
  • <图片名>_ocr.md:恢复出的 Markdown(需另加--recovery_to_markdown=True,recovery_to_markdown.py);
  • 表格区域会额外输出*.xlsx,图片区域输出裁剪后的*.jpg。

恢复阶段对位置信息的依赖点:

  • 分栏识别:sorted_layout_boxes(recovery_to_doc.py)依据每个区域bbox的 x 坐标与页面中线的关系,把区域归为左栏/右栏/通栏,还原论文双栏排版;
  • 段落判断:check_merge_method(recovery_to_markdown.py)比较文本行首 x 坐标与整个文本区 bbox 左边界的差值是否大于行高,以此区分"首行缩进"与"末行不满"两种段落标志;
  • 阅读顺序:_filter_text_res通过整行框与版面框的相交判断(_has_intersection,见 predict_system.py),把 OCR 结果归入对应版面区域,保证输出顺序正确。

八、注意事项

  1. 模型下载:文中的 wget 链接为 PP-OCRv3 系列官方推理模型(Paddle 3.0 生态目录),若版本更新导致链接失效,可前往 PaddleOCR 官方模型库查找替代下载地址;
  2. 输入路径:命令中的./docs/ppstructure/images/table_1.png、./docs/table/2.png为示例图片路径,请替换为实际存在的本地图片;
  3. 字典与字体必须配套:英文识别配en_dict.txt+simfang.ttf,中文识别配ppocr_keys_v1.txt+chinese_cht.ttf,错配会导致乱码或绘制失败;
  4. 性能影响:return_word_box=True会增加少量后处理计算量(逐字框换算 + 可视化绘制),对整体推理耗时影响很小,但建议生产环境按需开启;
  5. 中文单字框为近似结果:中文没有空格分词,单字框基于平均字宽近似估计,在字体宽度差异较大的场景可能存在 ±半个字宽的偏差,这是当前实现的固有特性(见cal_ocr_word_box源码)。

九、总结

return_word_box是 PP-Structure 文档恢复管线中"从内容识别走向结构化输出"的关键开关。开启后,识别结果从"一行一个框"升级为"一行一框 + 行内逐词/逐字小框",配合版面分析、表格识别与恢复模块,即可把一张图片/PDF 还原为带精确空间坐标的分层结构化数据——既可直接生成可编辑的 docx/markdown,也可作为 LLM/RAG 系统的"带位置索引"输入。相关延伸可继续阅读仓库内的 PP-Structure 总览、快速开始 以及 恢复为 Word/Markdown 文档。

【免费下载链接】PaddleOCRTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.项目地址: https://gitcode.com/GitHub_Trending/pa/PaddleOCR

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询