☰
Ouroboros TraceGuard 观察协议:用 EventStore 证据证明“类型化证据 + Verifier PASS”验收链路仍然成立
2026/10/10 1:41:46 网站建设 项目流程
  • AI Agent
  • 人工智能
  • 代码智能体
  • Agent 编排
  • AI 评测
  • CLI
  • 开发工具

【免费下载链接】ouroboros

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

项目地址:https://gitcode.com/gh_mirrors/ouroboros13/ouroboros
点击查看免费下载

这篇指南讲解 Ouroboros AgentOS 中 #978 类型化证据就绪度(typed evidence readiness)的可复现观察回环:如何从一个干净的主分支克隆出发,用单原子 AC 控制种子运行ooo run,再直接查询 SQLite EventStore 中的execution.ac.typed_evidence.observed事件,逐字段验证“typed evidence present → schema-valid typed evidence → verifier ran → verifier PASS”这条验收链路没有被 legacy 自报(self-report)路径旁路。读完本文,你可以独立执行一次受控观察运行、解读事件 payload 中的九个关键字段,并按仓库约定的模板产出可归档的观察报告。

1. 协议定位:只做观察,不做授权

该协议定义于 docs/agentos/traceguard-observation-protocol.md,服务于 #978 / #961 两项类型化证据就绪工作。在 post-#1082 的更广泛正面信号与 #978 P5 移除之后,它被用作回归检查:证明 AC 验收仍然完整流经“类型化证据 + Verifier PASS”这一门控,而不是悄悄退回到旧证据形态。

协议是**观察型(observation-only)**的,它明确不授权以下行为:

  • 重新引入 legacy 自报验收(legacy self-report acceptance);
  • 在“类型化证据 + Verifier PASS”之外改变ooo run的证据语义;
  • 新增任何 AgentOS substrate 表面;
  • 把单次受控运行当作未来发布就绪的依据。

协议的目标,是用 EventStore 证据来证明:原子 AC 的验收仍然以如下的完整链条完成——

typed evidence present -> schema-valid typed evidence -> verifier ran -> verifier PASS

这条链条上的每一环,都对应事件 payload 中一个可查询的布尔字段,这是后文所有查询与判定的基础。

2. 事件从哪来:源码层面的证据链

在跑命令之前,先理解这些字段是怎么产生的,判读结果时才不会迷路。

2.1 事件发射入口

execution.ac.typed_evidence.observed事件由执行器在原子 AC 完成时发射。发射入口是 execution_event_emitter.py 中的emit_atomic_typed_evidence_observed:它把data载荷原样包进BaseEvent,以aggregate_type=execution聚合身份追加到 EventStore。

2.2 payload 字段如何组装

载荷的组装逻辑在 parallel_executor.py 的_emit_atomic_typed_evidence_event中。观察协议查询的九个关键字段在这里逐一落值:

  • typed_evidence_present:等于typed_evidence is not None,即叶子是否真的产出了结构化证据记录;
  • typed_evidence_valid:等于typed_validation.ok,即证据记录是否通过了 profile 要求的 schema 校验(必填字段齐全、未命中rejected_if);
  • typed_evidence_error:提取/校验阶段的错误文本;
  • verifier_ran:等于verifier_verdict is not None,即验证器是否被实际调用并给出裁决;
  • verifier_passed:裁决是否通过;
  • verifier_failure_class/verifier_reasons/verifier_status:裁决未通过时的机读分类、原因列表与状态;
  • fat_harness_mode:运行是否处于“胖 harness”强制模式(构造ParallelExecutor时由fat_harness_mode参数注入,见 parallel_executor.py);
  • enforced:fat_harness_mode为真且验证器状态不是UNAVAILABLE时为真——验证器不可用(如 transcript 丢失)时会回退为仅观察,事件里可见enforced=false;
  • observe_only:not fat_harness_mode,即该次运行只做观察、不强制门控。

此外 payload 还携带enforcement_error(强制门被拒的原因)、required_fields(该 AC 生效证据 schema 的必填字段)、missing_fields/rejected_by/blocker(校验失败细节)以及typed_evidence_fields(实际提交的证据字段名)等诊断字段。

2.3 证据记录与校验结果的数据结构

typed_evidence与typed_validation两个核心对象定义在 evidence_schema.py:

  • EvidenceRecord:容器化的叶子证据字典,刻意保持宽容(保留原始映射与来源文本),schema 强校验交给验证器完成;
  • ValidationResult:校验结果,ok当且仅当没有缺失必填字段且没有任何rejected_if表达式命中;blocker是机读终端阻塞(如MISSING_AUTHORITY、MISSING_ACCESS、EXTERNAL_DEPENDENCY等),语义上“不是缺失证据”,而是合法前置条件未满足。

理解ValidationResult.blocker与missing_fields的区分很重要:它直接决定了负向结果判读时,你面对的是“叶子没交出证据”还是“叶子因环境限制无法交出证据”。

2.4 事件存储路径解析

观察脚本查询的 SQLite 库由 config/models.py 中的resolve_event_store_path解析:默认是配置目录下的ouroboros.db;如果配置文件中设置了persistence.database_path,则以该配置路径为准(相对路径相对配置文件所在目录解析;当配置路径与遗留ouroboros.db同时存在时,遗留路径优先)。这与观察协议脚本直接import ouroboros.config.models取路径的行为一致。

3. 提交前预检:从干净克隆出发

观察运行必须来自干净克隆,而不是脏的开发 checkout,否则无法把结果归因到某个 commit:

rm -rf /tmp/ouroboros-observation-main git clone <官方仓库地址> /tmp/ouroboros-observation-main cd /tmp/ouroboros-observation-main git checkout main git pull --ff-only git log --oneline -8

(原文档以项目官方仓库地址作为克隆源,此处按链接规范隐去具体 URL。)

对于 post-#1026 的观察,main分支必须包含 PR #1026 的合并提交。若缺少所需的类型化证据/验证器提交,本次运行仅具有诊断价值,不能记录为 clean-main 证据门就绪证据。git log --oneline -8就是用来人工确认合并提交的。

4. 受控种子:非重叠的单 AC 正向路径

种子设计的核心思路是:用一个原子 AC 同时创建实现和它的测试。早期两 AC 种子存在已知重叠——AC1 创建了test_hello.py之后,AC2 无法为同一文件产出新鲜的files_touched证据(files_touched是证据 schema 中的标准字段,提示词构建逻辑见 atomic_prompt_builder.py)。单 AC 设计消除了这个重叠源。

控制种子controlled-hello-seed-978.yaml全文如下:

goal: "Create a minimal Python hello function with a pytest verification." constraints: - "Use Python only." - "Do not add external dependencies beyond pytest." - "Keep the implementation minimal." acceptance_criteria: - | Create hello.py with a hello() function that returns exactly 'hello'. Create test_hello.py with a pytest test proving hello() returns exactly 'hello'. Run python -m pytest test_hello.py successfully. ontology_schema: name: "ControlledHello" description: "Minimal controlled seed for #978 typed evidence observation." fields: - name: "hello_function" field_type: "function" description: "hello() returns exactly hello." metadata: seed_id: "controlled_hello_978_single_ac" ambiguity_score: 0.0

各字段与仓库种子模型的对应关系,可参考 core/seed.py:

字段作用备注
goal单句任务目标明确交付物是“hello 函数 + pytest 验证”
constraints约束列表限 Python、限 pytest、限最小实现,压低变量
acceptance_criteria原子验收标准(本例仅 1 条)同时覆盖实现文件与测试文件,且以python -m pytest成功收尾
ontology_schema概念透镜,保证工作流语义连贯定义一个function类型字段hello_function
metadata.seed_id种子唯一标识便于在事件/日志中回溯本种子
metadata.ambiguity_score生成时模糊度评分模型默认 0.15,取值范围 0.0–1.0;本种子显式置 0.0,表示零歧义

种子刻意保持“可被测试独立复验”:即使不信任 harness 的裁决,你也可以在目标仓库里手动跑python -m pytest test_hello.py做 sanity check——这一点在第 9 节的负向判读中会用到。

5. 运行观察:完整命令

rm -rf /tmp/character-chat-978-observation mkdir -p /tmp/character-chat-978-observation cd /tmp/character-chat-978-observation git init cat > controlled-hello-seed-978.yaml <<'YAML' goal: "Create a minimal Python hello function with a pytest verification." constraints: - "Use Python only." - "Do not add external dependencies beyond pytest." - "Keep the implementation minimal." acceptance_criteria: - | Create hello.py with a hello() function that returns exactly 'hello'. Create test_hello.py with a pytest test proving hello() returns exactly 'hello'. Run python -m pytest test_hello.py successfully. ontology_schema: name: "ControlledHello" description: "Minimal controlled seed for #978 typed evidence observation." fields: - name: "hello_function" field_type: "function" description: "hello() returns exactly hello." metadata: seed_id: "controlled_hello_978_single_ac" ambiguity_score: 0.0 YAML PYTHONPATH=/tmp/ouroboros-observation-main/src \ uv --project /tmp/ouroboros-observation-main run ouroboros run controlled-hello-seed-978.yaml \ 2>&1 | tee observation-controlled-run-post-1026.log

要点:

  • 目标仓库是独立的/tmp/character-chat-978-observation空 git 仓库,与 ouroboros 源码克隆完全隔离;
  • 通过PYTHONPATH指向克隆的src/,再用uv --project ... run ouroboros run <seed>以该克隆的依赖环境运行 CLI,保证“被观察对象”就是那个 commit 的代码;
  • 全量 stdout/stderr 被tee到observation-controlled-run-post-1026.log,供第 7 节的 legacy 信号检查使用。

6. EventStore 汇总查询

运行结束后,用如下脚本直接查 EventStore 中最新的执行记录。脚本按execution_id分组、取最近一次执行,统计九个观察字段的取值分布:

python3 - <<'PY' import json import sqlite3 from collections import Counter, defaultdict from ouroboros.config.models import resolve_event_store_path db = resolve_event_store_path() con = sqlite3.connect(db) rows = con.execute( """ select timestamp, payload from events where event_type = 'execution.ac.typed_evidence.observed' order by timestamp desc """ ).fetchall() by_exec = defaultdict(list) for ts, payload in rows: data = json.loads(payload) if isinstance(payload, str) else payload execution_id = data.get("execution_id") or "unknown" by_exec[execution_id].append((ts, data)) latest_exec = next(iter(by_exec), None) print("latest_execution_id:", latest_exec) events = by_exec[latest_exec] print("typed_evidence_event_count:", len(events)) for key in [ "enforced", "fat_harness_mode", "typed_evidence_present", "typed_evidence_valid", "typed_evidence_error", "verifier_ran", "verifier_passed", "verifier_failure_class", "enforcement_error", ]: counts = Counter(str(data.get(key)) for _, data in events) print(key, dict(counts)) PY

判读口径与源码组装逻辑(第 2.2 节)一一对应:

  • typed_evidence_present/typed_evidence_valid同时为True,说明第一、二环(证据产出、schema 校验)通过;
  • verifier_ran为True,说明验证器被实际调用;verifier_passed为True说明第四环通过;
  • enforced为True表示该次运行确实以强制门(而非仅观察)方式执行;若看到enforced=False且fat_harness_mode=True,通常意味着验证器状态为UNAVAILABLE(transcript 丢失或仅有不可重放内容),需要另行排查;
  • typed_evidence_error、verifier_failure_class、enforcement_error三个字段是负向定位的第一手线索。

7. Legacy 回退信号检查

grep -iE "legacy|self.report|self-report|self_report|fallback" \ /tmp/character-chat-978-observation/observation-controlled-run-post-1026.log || true

这一步只是日志信号检查:grep 无命中本身不能证明所有内部回退分支都不可达。

从源码结构看,legacy 路径确实仍存在于执行器中,但已被收窄:parallel_executor.py 中,当fat_harness_error非空且运行启用了check_package_gate时,该拒绝会被改写为legacy_rejection并清空fat_harness_error——也就是说 legacy 自报语义只在 check package 门语境下作为历史兼容分支存在,正常受控运行中验收拒绝仍走类型化证据 + 验证器裁决路径,并以parallel_executor.ac.verifier_rejected结构化日志暴露完整原因。这解释了为什么协议要求“证据门语义”以 EventStore 事件为准、以日志 grep 仅作辅证。

8. 判定标准:正信号、完整通过与负向结果

8.1 正向受控信号(positive controlled signal)

一次受控运行构成正信号,当且仅当至少有一个被观察到的 AC 满足:

  • typed_evidence_present=true
  • typed_evidence_valid=true
  • verifier_ran=true
  • verifier_passed=true
  • 日志中没有可见的 legacy 自报回退验收信号

一次**完整受控通过(full controlled pass)**更强,还应显示 CLI 运行本身成功。

8.2 负向 / 阻塞结果判读

观测结论
typed_evidence_present=false或typed_evidence_valid=falseprompt / extractor / schema 接缝仍然被阻塞——叶子没交出证据,或交出的证据不满足 schema
verifier_ran=false验证器调用在裁决之前仍被门住
verifier_passed=false检查verifier_reasons/enforcement_error;通常是证据匹配问题或真实失败
手动 pytest 通过但 verifier 失败实现可能正确,但验收没有走通证据门;按协议应视为回归阻塞项,而不是可忽略的偏差

最后一行是本协议最具操作性的判据:它强制你把“代码对了”和“证据门放行了”当作两个独立命题。实现正确但门控未放行,说明回归发生在证据链上,必须按 #978 回归处理。

9. 报告模板

结果需发布到 #978;若改变了 SSOT,还需在 #961 上同步状态。仓库约定的报告模板如下:

## #978 Observation batch N — post-#1026 clean-main controlled run Date: YYYY-MM-DD TZ Ouroboros commit: <sha> Target repo: /tmp/character-chat-978-observation Seed: controlled-hello-seed-978.yaml Run command: bash PYTHONPATH=/tmp/ouroboros-observation-main/src uv --project /tmp/ouroboros-observation-main run ouroboros run controlled-hello-seed-978.yaml CLI result: - Total ACs: - Succeeded: - Failed: - Skipped: - Exit status: Typed evidence: - execution id: - event count: - typed_evidence_present=true: - typed_evidence_valid=true: - verifier_ran=true: - verifier_passed=true: - typed_evidence_error values: - enforcement_error values: Legacy fallback: - Used / not observed / unknown: - Evidence: Manual sanity check: - Files created: - `python -m pytest test_hello.py` result: Conclusion: - Positive controlled signal: yes/no - #978 P5 regression suspected: yes/no - Next observation/fix needed:

模板中“Typed evidence”一栏对应第 6 节脚本的九个输出,“Legacy fallback”一栏对应第 7 节 grep 的结果,“Manual sanity check”一栏则是独立于 harness 的交叉验证——三者共同保证报告同时覆盖了事件证据、日志信号与手工复验三个层面。

10. Post-P5 就绪边界

单次受控运行不足以支撑未来发布信心。post-#1082 的更广泛观察为 #978 P5 移除提供了正信号;在此之后,任何失败都应被当作证据门回归或后续修复项对待,而不是恢复 legacy 自报验收的理由。这条边界与第 1 节“协议不做什么”的约束互相咬合:观察协议的存在价值恰恰是——用可复现的 EventStore 证据持续证明证据门仍然有效,并让任何旁路回归第一时间显形。

与本协议相邻、可进一步深入的材料包括同一目录下的 traceguard-vs-legacy-benchmark.md(TraceGuard 与 legacy 路径的基准对比)和 fat-harness-baseline-metrics.md(胖 harness 基线指标),它们与本文的字段判读口径共同构成 #978 证据就绪度的完整文档面。

  • AI Agent
  • 人工智能
  • 代码智能体
  • Agent 编排
  • AI 评测
  • CLI
  • 开发工具

【免费下载链接】ouroboros

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

项目地址:https://gitcode.com/gh_mirrors/ouroboros13/ouroboros
点击查看免费下载

相关推荐

上一篇:RobotGo快速入门:30分钟上手Go语言桌面自动化脚本开发
下一篇:zfile的容器化最佳实践:镜像优化与资源限制

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询