☰
Harbor 中的 DABstep 金融数据分析基准适配:instruction.md 提示模板与评测链路全解析
2026/10/11 14:10:03 网站建设 项目流程

【免费下载链接】harbor

Framework for evaluating and improving agents

项目地址:https://gitcode.com/gh_mirrors/harbor17/harbor
点击查看免费下载

DABstep(Data Agent Benchmark for Multi-step Reasoning)是 Adyen 发布的金融数据分析 Agent 评测基准:Agent 拿到支付交易数据后,必须编写 Python 代码完成需要多步推理的业务问题。本文围绕 Harbor 仓库中 DABstep 适配器的instruction.md任务提示模板,从模板语义、渲染机制、任务目录结构、环境构建、评分器到 parity 验证逐层拆解,读完即可理解该适配器如何把 HuggingFace 上的 460 个原始任务转换为可在 Harbor 中运行、评分与复现的标准化任务,并掌握用uv run harbor jobs start/uv run harbor trials start运行 DABstep 评测的完整实操。

一、DABstep 基准:用支付数据考验多步推理

DABstep 是 Adyen 面向金融数据分析场景的 Agent 评测基准,核心评估方式是:向 Agent 提供支付交易记录、费用表、商户数据等真实结构的数据文件,要求 Agent 编写 Python 代码回答诸如「德国境内 Visa 交易的平均交换费是多少」这类业务问题。这类查询通常需要多数据源关联(join)、业务规则套用与聚合计算,这正是 DABstep 强调的多步推理能力(见 适配器 README)。

该适配器在 Harbor 中的定位与规模(以仓库内 adapter_metadata.json 与 README 为准):

  • 任务类型:金融数据分析 + Python 代码生成;
  • 数据集规模:460 个任务,去重后 454 个唯一任务;
  • 默认切分(default split):450 个任务(72 easy + 378 hard),答案从公开的task_scores榜单提交中提取;
  • 开发切分(dev split):10 个任务,数据集直接内置标准答案;
  • License:CC-BY-4.0;
  • Oracle 验证:454 个唯一任务全部通过 oracle 测试,reward 100%。

二、instruction.md 模板逐段解析:Agent 的任务指令

adapters/dabstep/template/instruction.md是每个 DABstep 任务的提示词模板,它在生成任务时被{question}与{guidelines}两个占位符替换。模板全文仅 11 行,却定义了完整的任务契约:

You are an expert data analyst and you will answer factoid questions by referencing files in the data directory: `/app/data/` Don't forget to reference any documentation in the data dir before answering a question. Here is the question you need to answer: {question} Here are the guidelines you MUST follow when answering the question above: {guidelines} Before answering the question, reference any documentation in the data dir and leverage its information in your reasoning / planning. When you have computed the final answer, write ONLY the final answer to `/app/answer.txt` (e.g. if the answer is 42, the file should contain just `42`).

逐句拆解其语义:

  1. 角色与数据入口(第 1 行):Agent 被设定为「专家数据分析师」,必须通过/app/data/目录中的文件回答事实型问题。这意味着 Agent 的第一步是勘察数据目录,而不是凭常识作答。
  2. 强制查阅文档(第 2、8 行):模板两次强调「回答问题前先参考数据目录中的任何文档」。对应 Dockerfile 中预下载的manual.md、payments-readme.md——数据目录里不仅有原始数据,还有业务手册与字段说明,Agent 必须从中获取收费规则、字段口径等信息才能正确计算。
  3. 问题与准则占位符(第 4、6 行):{question}与{guidelines}在生成任务时由真实任务内容替换(见第三节的渲染机制)。
  4. 答案输出协议(第 11 行):只把最终答案写入/app/answer.txt,并以「如果答案是 42,文件里就只写 42」为例说明格式要求——答案必须干净、无解释性文字,这是后续模糊评分器能正确比对的前提。

这条输出协议是 Harbor 适配相对原始评测框架的关键改造点:README 明确说明「Prompt adapted from the original evaluation harness to instruct agents to write final answers to/app/answer.txt」,即原始 DABstep 评测的提示被调整为让 Agent 将答案落到固定文件,从而与 Harbor 的 verifier 机制(test.sh读取该文件)对接。

三、模板渲染机制:从模板到真实任务

instruction.md不是直接使用,而是由适配器代码在生成任务目录时完成占位符替换。adapter.py 中的DABstepTask(L23-L40)将数据集每条记录映射为任务对象:

class DABstepTask: """Represents a single DABstep task.""" def __init__(self, record: dict, answer: str): self.task_id = str(record["task_id"]) self.question = record["question"] self.guidelines = record["guidelines"] self.level = record["level"] self.answer = answer @property def id(self) -> str: return f"dabstep-{self.task_id}" @property def difficulty(self) -> str: return "easy" if self.level == "easy" else "hard"

DABstepAdapter._prepare_task(L66-L105)负责渲染整棵任务目录树,其中与 instruction.md 相关的部分为:

# --- instruction.md --- instruction = (TEMPLATE_DIR / "instruction.md").read_text() instruction = instruction.replace("{question}", task.question) instruction = instruction.replace("{guidelines}", task.guidelines) (output_dir / "instruction.md").write_text(instruction)

同一方法内还对 task.toml 替换{difficulty}与{tags}、对solution/solve.sh与tests/test.sh替换{answer},并把适配器根目录的scorer.py复制到任务的tests/下。也就是说:一份模板 + 逐任务真实数据 = 一个可直接运行的 Harbor 任务。

四、生成的任务目录结构与 Harbor 约定

README 给出每个任务目录的最终形态:

dabstep-{task_id}/ ├── task.toml ├── instruction.md ├── environment/ │ └── Dockerfile ├── solution/ │ └── solve.sh └── tests/ ├── test.sh └── scorer.py

这一结构与 Harbor 核心的任务路径模型完全一致。src/harbor/models/task/paths.py(L13-L31)的TaskPaths类注释定义了同样的约定:instruction.md位于任务根目录,solution/会被复制进容器/solution(由 OracleAgent 使用),tests/会被复制进容器/tests(由 Evaluator 使用),脚本按.sh>.bat优先级发现。DABstep 适配器生成的solve.sh、test.sh正是这两个入口。

其中 task.toml 是 Harbor 任务元数据与资源配额的声明文件:

version = "1.0" [metadata] author_name = "Adyen" author_email = "unknown" difficulty = "{difficulty}" category = "data-analysis" tags = [{tags}] [verifier] timeout_sec = 600.0 [agent] timeout_sec = 1800.0 [environment] build_timeout_sec = 600.0 cpus = 1 memory = "4G" storage = "8G"
  • difficulty由适配器按level字段换算为easy/hard;tags由适配器拼为"dabstep", "data-analysis", "financial", "{level}"(见 adapter.py);
  • [verifier]与[agent]分别限定评分与 Agent 运行的超时(600 秒 / 1800 秒);
  • [environment]声明容器配额:1 CPU、4G 内存、8G 存储、600 秒构建超时。

五、运行环境:预置数据文件的 Docker 镜像

environment/Dockerfile 从ghcr.io/laude-institute/t-bench/ubuntu-24-04:20250624基础镜像构建,安装pandas后,在构建阶段就把 DABstep 的 7 个共享数据文件(约 24.2 MB)从 HuggingFace 下载进镜像的/app/data/:

文件用途
payments.csv支付交易记录(主数据源)
fees.json费用表
manual.md业务规则手册(Agent 必须参考)
merchant_data.json商户数据
merchant_category_codes.csv商户类别代码
payments-readme.md支付数据字段说明文档
acquirer_countries.csv收单国家数据

「预下载进镜像」是 README 明确记录的改造点:Agent 运行期间无需联网即可访问/app/data/。注意一个基础前置条件——docker build阶段需要联网,因为数据文件是在构建时下载的;此外评分环境依赖datasets、pandas、pyarrow等包(见 README 的 Infrastructure Requirements)。

六、评分链路:test.sh + scorer.py 的模糊匹配

评测的判定闭环由 tests/test.sh 与根目录的 scorer.py 组成。

test.sh的逻辑:

  1. 先检查/app/answer.txt是否存在,不存在直接写0到/logs/verifier/reward.txt;
  2. 读取 Agent 答案首行(head -1)作为待评答案;
  3. 读取适配器注入的期望答案({answer}占位符被替换);
  4. 调用python3 /tests/scorer.py "$agent_answer" "$expected_answer";
  5. 按退出码把1(正确)或0(错误)写入/logs/verifier/reward.txt。

/logs/verifier/reward.txt是 Harbor verifier 的标准奖励输出约定——仓库中多处任务模板与adapter_review.py检查项都以此为准(例如 src/harbor/cli/adapter_review.py 的检查项「tests/test.sh writes reward to /logs/verifier/reward.txt」)。

scorer.py是从官方 DABstep 评分器移植的自定义实现(文件头注明来源与 CC-BY-4.0 License),核心入口是question_scorer(L88-L103),比较策略按答案形态分层:

def question_scorer(input1: str, input2: str) -> bool: input1 = input1.strip().lower() input2 = input2.strip().lower() if is_numeric_with_commas(input1) or is_numeric_with_commas(input2): num1 = extract_numeric(input1) num2 = extract_numeric(input2) if num1 is not None and num2 is not None: return compare_numeric(num1, num2) return False if ";" in input1 or ";" in input2 or "," in input1 or "," in input2: return compare_lists(input1, input2) num1 = extract_numeric(input1) num2 = extract_numeric(input2) if num1 is not None and num2 is not None: return compare_numeric(num1, num2) return compare_strings(input1, input2)

评分器的容忍细节(scorer.py):

  • 数值比较(compare_numeric,L43-L55):先做小数位数对齐的四舍五入比较,再用math.isclose以rel_tol=1e-4, abs_tol=1e-4做浮点容差;小于 1 的值直接按该容差判定。因此「42」与「42.0001」这类微小偏差不会误判为错。
  • 字符串比较(compare_strings,L58-L68):先去除非单词字符后精确匹配;单侧单词数为 1 时允许子集包含;否则用SequenceMatcher相似度 > 0.95 判定。这解释了 README 中「Fuzzy matching with numeric tolerance (rel_tol=1e-4) and string similarity (SequenceMatcher > 0.95)」的表述。
  • 列表比较(compare_lists,L71-L85):支持[...]包裹、[,;]分隔的答案,排序后逐元素递归评分。
  • 数值形态识别(is_numeric_with_commas,L18-L28):支持带$前缀、千位逗号(如1,234.56)等金融数据常见写法,extract_numeric(L31-L40)会剥掉,与$再提取数值。

值得注意的边界:README 记录有 3 个任务(ID:2566、2522、2521)的标准答案就是空字符串,这是 DABstep 中的合法答案,评分器可以正确处理。

七、数据切分与答案提取策略

DABstep 的两个切分 schema 一致(task_id、question、guidelines、level、answer),关键差异在答案来源:

  • default split(450 任务):数据集本身不含标准答案,run_adapter.py 的extract_answers_from_task_scores(L56-L100)从公开的task_scoresparquet 文件下载榜单提交,筛选score == True的正确提交,每个任务取最短的正确答案(最可能是干净值而非冗长回复),再经clean_answer(L25-L53)后处理:
    • 剥离 markdown 代码块(json .../...);
    • 剥离单元素列表包装(["B"]→B);
    • 多行答案只取第一个非空行(丢弃 Agent 追加的解释性文字)。
  • dev split(10 任务):数据集直接带answer字段,run_adapter.py直接从行记录读取。

6 个任务 ID(5、49、70、1305、1681、1753)在两个切分间重叠,内容与答案一致。

八、实操:生成任务与运行评测

8.1 生成任务目录

在adapters/dabstep目录下用uv run执行(也可通过--output-dir指定输出位置):

# 生成 default split(450 任务,答案自动从 task_scores 提取) uv run run_adapter.py # 生成 dev split(10 任务,内置标准答案) uv run run_adapter.py --split dev # 使用预提取的答案文件 uv run run_adapter.py --answers-file /path/to/answers.json # 自定义输出目录 uv run run_adapter.py --output-dir ../../datasets/dabstep

CLI 参数由 run_adapter.py 定义:--split限定default/dev,--answers-file可跳过联网提取答案的环节。

8.2 通过数据集注册表运行作业

# Oracle agent(参考解,验证任务可解性) uv run harbor jobs start -d dabstep # 指定 Agent 与模型 uv run harbor jobs start -d dabstep -a terminus-2 -m "anthropic/claude-haiku-4-5"

8.3 通过作业配置文件运行

adapters/dabstep/dabstep.yaml 是仓库内现成的作业配置:

jobs_dir: jobs n_attempts: 1 timeout_multiplier: 1.0 orchestrator: type: local n_concurrent_trials: 1 quiet: false environment: type: docker force_build: true delete: true env: - OPENAI_API_KEY=${OPENAI_API_KEY} - GEMINI_API_KEY=${GEMINI_API_KEY} - ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY} agents: - name: terminus-2 model_name: anthropic/claude-haiku-4-5 datasets: - path: datasets/dabstep
uv run harbor jobs start -c adapters/dabstep/dabstep.yaml -a terminus-2 -m "anthropic/claude-haiku-4-5"

配置要点:environment.delete: true表示每次试用后删除容器,force_build: true强制重建镜像(配合 Dockerfile 构建期下载数据);模型名以anthropic/为前缀的完整 provider 路径形式给出。

8.4 运行单个试用(trial)

# 单个任务 + oracle uv run harbor trials start -p datasets/dabstep/dabstep-5 # 单个任务 + 指定 Agent 与模型 uv run harbor trials start -p datasets/dabstep/dabstep-5 -a terminus-2 -m "anthropic/claude-haiku-4-5"

-p/--path指向任务目录,-a/--agent与-m/--model选择 Agent 和模型(参数定义见 src/harbor/cli/trials.py),另有--trials-dir、--timeout-multiplier、--trial-name等可选参数。

8.5 环境前置条件

  • 已安装并运行 Docker;
  • 从仓库根目录执行uv sync --extra dev安装 Harbor;
  • Python 依赖:pip install datasets pandas pyarrow;
  • 数据集公开无门槛,无需 API key(构建镜像阶段需联网下载数据文件)。

九、Parity 验证:与原始基准的对齐证据

DABstep 适配器通过分层抽样 + 双端跑分验证了与原始基准的一致性。

抽样方法(generate_parity_sample.py):从 default split 按难度分层随机抽样,seed=42,目标 130 任务——easy 21(16%)、hard 109(84%),产出 parity_sample_task_ids.txt(文件头注释记录了抽样分布与用法)。

跑分设置与结果(parity_experiment.json 与 README 表格):Harbor 侧用terminus-2+claude-haiku-4-5(Docker 环境),fork 侧用claude -pCLI,各跑 4 个 trial:

Agent模型Metric样本量原始基准Harbor 适配
claude-codeclaude-haiku-4-5Accuracy (%)130 任务(占全集 28.9%)37.69 ± 0.4436.92 ± 0.31
claude-codeclaude-haiku-4-5Accuracy (%) — easy21 个 easy 任务84.52 ± 1.1984.52 ± 1.19
claude-codeclaude-haiku-4-5Accuracy (%) — hard109 个 hard 任务28.67 ± 0.6927.75 ± 0.44

结论(README 明确表述):easy 任务 4 个 trial 全部精确对齐(0 pp 差异),总体差异全部来自 hard 任务(0.92 pp,处于统计波动范围内)。

十、注意事项与验证状态汇总

  • 答案来源按切分不同:default split 的答案是提取的(最小正确提交 + 清洗),dev split 的答案直接内置;若task_scores中某任务无score=True提交,则该任务会被跳过(DABstepAdapter.__init__会记录No answer found for task_id=...警告)。
  • 构建期联网:docker build时必须能访问 HuggingFace 下载 7 个共享数据文件。
  • 验证状态:Oracle 验证覆盖全部 454 个去重任务(100% reward);parity 验证基于 130 任务 × 4 trial(<1% 差异)。以上均为仓库内 README 与 parity_experiment.json 记录的既有事实,新场景下的表现需自行实测。
  • 难度分布:default split 中 easy 72/450(16%)、hard 378/450(84%),dev split 分布一致,便于按需选择快速验证(dev)或全量评测(default)。

十一、总结

adapters/dabstep/template/instruction.md是整个 DABstep 适配器的「任务契约原点」:它定义了 Agent 的角色(数据分析专家)、数据入口(/app/data/)、强制查阅文档的行为、以及唯一答案出口(/app/answer.txt)。围绕这条模板,适配器在 Harbor 中串联起完整的评测链路——run_adapter.py负责答案提取与渲染,Dockerfile 预置数据,test.sh+scorer.py实现带数值容差与字符串相似度的模糊评分,dabstep.yaml与 CLI 命令提供一键运行入口,parity_experiment.json则记录了与原基准的对齐证据。对希望把「数据密集型、多步推理、自由答案格式」类基准接入 Harbor 的开发者而言,DABstep 适配器是一个结构清晰、可直接复用的参考实现。

【免费下载链接】harbor

Framework for evaluating and improving agents

项目地址:https://gitcode.com/gh_mirrors/harbor17/harbor
点击查看免费下载

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询