☰
agno 文档抽取实战:用 PDF 文件输入与 Pydantic 输出模式构建结构化文档标注
2026/9/29 14:58:18 网站建设 项目流程

agno 文档抽取实战:用 PDF 文件输入与 Pydantic 输出模式构建结构化文档标注

【免费下载链接】agnoBuild, run, and manage agent platforms.项目地址: https://gitcode.com/GitHub_Trending/ag/agno

导读

本文基于 agno 仓库中cookbook/data_labeling/_16_document_extraction/目录下的三个示例与对应测试记录,完整讲解如何将多页 PDF 输入给 Agent,并通过output_schema把抽取结果约束为类型化的 Pydantic 对象。你将掌握三种递增的抽取形态:顶层文档元数据抽取、嵌套行项目(line item)抽取、以及带字段级置信度的抽取,并了解它们在发票字段落库、合同条款审查队列、PDF 语料结构化索引等生产标注场景中的落点。

一、示例概览与测试结论

cookbook/data_labeling/_16_document_extraction/目录以一份公开的泰式菜谱 PDF(ThaiRecipes.pdf)作为演示输入,包含三个脚本:

  • basic.py:抽取文档级元数据(标题、菜系、语言、菜谱数量)。
  • with_line_items.py:在元数据之外抽取嵌套的List[Recipe],即“行项目”形态。
  • with_confidence.py:为每个字段额外输出high / medium / low置信度。

目录下的 TEST_LOG.md 记录了 2026-07-18 在 agno 2.7.4、模型gemini-3.5-flash上的真实验证结果,三个脚本全部PASS:

脚本核心能力实测结果(摘录)
basic.pyPDF 文件输入 + 类型化输出返回RecipeBook(title='Thai SELECT Cookbook', cuisine='Thai', language='English', recipe_count=10),单次模型调用,7449 tokens,约 3.9s
with_confidence.py嵌套模型 +Literal枚举四个字段均返回 confidence "high",约 6.0s
with_line_items.py嵌套List行项目抽取返回 10 个菜谱(含 Pad Thai Goong Sod、Tom Kha Gai、Gluai Buat Chi 等),course字段全为 null(文档未标注菜系类别,模型如实留空),约 9.4s

这几个测试结果不仅证明示例可运行,也印证了关键工程行为:文档里没有的信息,模型会按指令返回 null,而不是臆造。

二、基础形态:抽取顶层文档元数据

2.1 定义输出 Schema

basic.py的核心是用 Pydantic 模型声明“从文档里要哪些字段”。每个字段的description会作为结构化输出提示的一部分交给模型,因此描述要写得具体、可执行:

from typing import Optional from pydantic import BaseModel, Field class RecipeBook(BaseModel): title: Optional[str] = Field(None, description="Book or document title") cuisine: Optional[str] = Field(None, description="Cuisine or culinary tradition") language: Optional[str] = Field(None, description="Language of the document") recipe_count: Optional[int] = Field( None, description="Number of distinct recipes in the document" )

字段全部为Optional并默认None,配合指令“If a field is not present, leave it null”,让模型在信息缺失时安全地输出空值。

2.2 创建 Agent 并传入 PDF

from agno.agent import Agent, RunOutput from agno.media import File instructions = """\ Extract document-level metadata from the attached PDF. Use exactly what the document shows. If a field is not present, leave it null. """ agent = Agent( model="google:gemini-3.5-flash", instructions=instructions, output_schema=RecipeBook, ) if __name__ == "__main__": url = "https://agno-public.s3.amazonaws.com/recipes/ThaiRecipes.pdf" run: RunOutput = agent.run("Extract document metadata.", files=[File(url=url)]) pprint({"url": url, "result": run.content})

关键点在于files=[File(url=url)]:File是 agno 媒体输入模型,声明于 libs/agno/agno/media/media.py,支持id、url、filepath、content(原始字节)、external、media_reference六种来源,并要求至少提供其一(由check_at_least_one_source校验器强制)。也就是说,既可以用 URL 传入公网 PDF,也可以把本地发票/合同 PDF 换成File(filepath="..."),或者直接传入字节流content。

同时File.mime_type有白名单校验(见 media.py),允许application/pdf、application/json、.docx/.xlsx/.pptx等 Office Open XML 格式以及text/plain等文档类 MIME 类型。

2.3 输出模式的底层机制

output_schema定义于 libs/agno/agno/agent/agent.py:它接受一个 Pydantic 模型类或符合供应商期望的 JSON schema 字典,Agent 会把 schema 注入请求,将模型响应约束为对应 JSON,随后parse_response=True(默认开启)会把 JSON 反序列化回 Pydantic 对象,最终通过RunOutput.content直接拿到类型化结果。若关闭解析,返回的将是 JSON 字符串。此外还支持parser_model(用第二模型兜底解析)、structured_outputs(供应商强制结构化输出开关)与use_json_mode(把 schema 的 JSON 描述写入系统消息)等进阶配置。

三、行项目形态:抽取嵌套子对象列表

生产标注中最常见的形态不是一层元数据,而是“一条单据 + 多行明细”,例如发票行项目、对账单交易、合同条款。with_line_items.py演示了这种形态:外层是文档元数据,内层是List[Recipe]。

from typing import List, Optional from pydantic import BaseModel, Field class Recipe(BaseModel): name: str = Field(..., description="Recipe name as printed") course: Optional[str] = Field(None, description="Appetizer, main, dessert, etc.") prep_time_minutes: Optional[int] = None class RecipeBook(BaseModel): title: Optional[str] = None cuisine: Optional[str] = None recipes: List[Recipe] = Field(default_factory=list)

注意Recipe.name是必填字段(...),其余为可选——行项目至少要保证“名称”可抽取;recipes用default_factory=list确保文档中没有任何明细时也能反序列化出空列表,而不是抛出校验错误。

指令层面强调“Do not invent recipes or paraphrase names”(不得发明菜谱或改写名称),这是行项目抽取中防止模型幻觉的关键提示:

instructions = """\ Extract document metadata and every distinct recipe from the attached PDF. Each recipe entry should reflect what the document actually shows; leave fields null if not present. Do not invent recipes or paraphrase names. """

测试日志给出的结果与指令一致:模型返回了 10 个菜谱,course全部为 null——因为菜谱 PDF 本身没有按“开胃菜/主菜/甜点”分类,模型没有自行脑补。这个行为验证了行项目抽取的可靠性边界:只抽文档中真实存在的内容,缺省字段留空。实测该形态约 9.4s,是三种示例中耗时最长、信息量最大的一种。

四、置信度形态:为每个字段附加 high / medium / low

当输入 PDF 质量参差不齐(扫描件、传真件、混合语言)时,下游需要把不确定的字段路由到人工复核。with_confidence.py用嵌套模型 +Literal枚举实现字段级置信度:

from typing import Literal, Optional from pydantic import BaseModel Confidence = Literal["high", "medium", "low"] class ConfidentField(BaseModel): value: Optional[str] = None confidence: Confidence class RecipeBook(BaseModel): title: ConfidentField cuisine: ConfidentField language: ConfidentField recipe_count: ConfidentField # 以字符串持有,便于对数量字段应用统一的置信度结构

指令中明确定义了三档置信度的语义,并要求保守:

instructions = """\ Extract document metadata. For each field, report confidence: - high - explicit in the document - medium - inferred from structure or context - low - guessed, partly obscured, or ambiguous Be conservative. Mark unsure fields low. """

这种“值 + 置信度”的嵌套结构直接服务于人工复核队列:只有confidence != "high"的字段进入人工队列,high字段直接落库。测试日志显示,对于这份清晰的菜谱 PDF,四个字段均返回 "high" 置信度,符合预期。把recipe_count特意声明为字符串类型持有,则避免了整数字段在“置信度 + 数值”结构中难以统一处理的问题——从源码注释可见这是有意为之。

五、生产落地与邻近方案

5.1 典型生产场景

README.md 明确了本示例在生产标注中的定位与适用场景:

  • 把发票 / 收据 / 对账单字段抽成数据库行;
  • 抽取合同条款进入审查队列;
  • 为 PDF 语料库建立结构化索引。

5.2 如何换成自有文档

示例使用公开菜谱 PDF 以便开箱即跑;生产使用时只需替换输入来源与 schema:

python cookbook/data_labeling/_16_document_extraction/basic.py python cookbook/data_labeling/_16_document_extraction/with_line_items.py python cookbook/data_labeling/_16_document_extraction/with_confidence.py
  • 把File(url=...)换成File(filepath="/path/to/invoice.pdf")或File(content=bytes);
  • 把RecipeBook/Recipe换成Invoice/InvoiceLineItem、Contract/ContractClause、LabReport等业务模型;
  • 将model换为当前环境可用的模型,并确保已配置对应的GOOGLE_API_KEY等凭据。

5.3 邻近能力对照

  • 如果只需要“文档属于哪个类型”这种单标签结果,应使用 _15_document_classification 而不是文档抽取;
  • 若需要在抽取之上叠加多 Agent 质量审查,参考 _18_quality_review;
  • 三种示例的 schema 能力(嵌套模型、Literal枚举、List字段)同样适用于 agno 其他结构化输出场景,相关配置参数(output_schema、parse_response、structured_outputs、use_json_mode)统一由 agent.py 提供。

六、小结

cookbook/data_labeling/_16_document_extraction/用一个菜谱 PDF 完整演示了 agno 文档抽取的三种递增形态:顶层元数据、嵌套行项目、字段级置信度。核心机制是File(URL / 本地路径 / 字节流)与output_schema(Pydantic 结构化输出)的组合;测试日志验证了模型在缺省字段上返回 null、在行项目上不臆造内容的可靠行为。对于发票落库、合同条款提取、PDF 语料索引等生产标注任务,这套模式可以直接迁移:换 PDF、换 schema、换模型即可。

【免费下载链接】agnoBuild, run, and manage agent platforms.项目地址: https://gitcode.com/GitHub_Trending/ag/agno

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询