☰
用 nf-test 为 Nextflow 与 nf-core 组件构建可靠测试:模块、子工作流与流水线的测试全指南
2026/10/10 14:25:20 网站建设 项目流程

用 nf-test 为 Nextflow 与 nf-core 组件构建可靠测试:模块、子工作流与流水线的测试全指南

【免费下载链接】scientific-agent-skillsTurn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 190,000+ scientists worldwide. 165 ready-to-use validated skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.项目地址: https://gitcode.com/GitHub_Trending/cl/scientific-agent-skills

nf-test 是 Nextflow 生态的标准测试框架,也是 nf-core 社区对每个模块(module)、子工作流(subworkflow)和流水线(pipeline)的强制测试要求。本文以 scientific-agent-skills 仓库中 Nextflow 技能 的测试参考文档为核心,系统讲解 nf-test 的安装初始化、测试文件结构、三种测试作用域(process/workflow/pipeline)、断言与快照测试的完整写法,并结合仓库内 developing.md 与 nf-core-tools.md 的源码级约定,给出可复制、可运行、可直接进入 CI 的实战方案。读完后你将能够独立为任意 Nextflow 模块编写并维护 nf-test 测试套件,并通过 nf-core 工具链完成测试、lint 与 CI 集成。

为什么是 nf-test:Nextflow 测试的行业标准

Nextflow 是一个用于构建可复现、可移植、可扩展数据流水线的工作流语言与运行时,在生物信息学领域占据主导地位,同时适用于任何数据密集型计算(详见 SKILL.md 的 Overview)。nf-core 则是建立在 Nextflow 之上的社区,负责维护生产级流水线、可复用模块以及nf-core工具链。

在这样一个以"复用与标准化"为核心的生态里,测试不再是可选项:nf-test 是 Nextflow 的标准测试框架,nf-core 要求每个模块、子工作流和流水线都必须具备 nf-test 覆盖。这意味着:

  • 模块作者提交代码前,必须提供可运行的 nf-test 测试(含 stub 测试)并提交快照(snapshot)文件;
  • nf-core modules lint会校验测试文件是否齐备;
  • CI 中通过--changed-since只对变更组件执行测试,保证每次合并的可靠性。

本指南对应的核心文档位于 skills/nextflow/references/testing.md,是 Nextflow 技能体系中负责测试的独立自洽章节。

安装与初始化(Setup)

nf-test 的安装方式有两种,任选其一:

# 方式一:通过 conda/bioconda 安装 conda install -c bioconda nf-test # 方式二:官方安装脚本(自包含) curl -fsSL https://get.nf-test.com | bash

安装完成后,在项目根目录执行初始化:

nf-test init # 创建 nf-test.config + tests/ 脚手架

nf-test init会在项目中生成两个关键产物:

  1. nf-test.config:集中设置测试目录、默认 profile 和插件;
  2. tests/脚手架:测试文件的组织目录。

测试文件的命名与位置遵循严格约定:

  • 测试文件以.nf.test结尾;
  • 测试文件放在被测组件旁边,例如tests/main.nf.test(在 nf-core 组件目录下则为modules/nf-core/<tool>/<subtool>/tests/main.nf.test);
  • 期望结果存储在同级的.nf.test.snap快照文件中,例如tests/main.nf.test.snap。

从 developing.md 的模块目录解剖可以看到 nf-core 的标准布局,测试是模块不可分割的一部分:

modules/nf-core/samtools/sort/ ├── environment.yml # Conda channels + 固定版本依赖 ├── main.nf # 被测 process ├── meta.yml # 记录 I/O 接口 + 工具信息(schema 校验) └── tests/ ├── main.nf.test # nf-test 测试(必需,且必须包含 stub 测试) └── main.nf.test.snap

测试文件结构:三种作用域

.nf.test文件用 Groovy 编写,外层作用域(scope)决定了测试对象,共有三种:

作用域测试对象说明
nextflow_process单个 process/模块最常用,直接驱动一个模块
nextflow_workflow子工作流测试由多个模块组成的 (sub)workflow
nextflow_pipeline整条流水线驱动整个main.nf

一个典型的nextflow_process测试文件结构如下(以 SAMTOOLS_SORT 为例):

nextflow_process { name "Test SAMTOOLS_SORT" script "../main.nf" // 被测模块 process "SAMTOOLS_SORT" tag "modules" tag "modules_nfcore" tag "samtools" tag "samtools/sort" test("sarscov2 - bam") { when { process { """ input[0] = [ [ id:'test', single_end:false ], file(params.modules_testdata_base_path + 'genomics/sarscov2/illumina/bam/test.bam', checkIfExists: true) ] """ } } then { assertAll( { assert process.success }, { assert snapshot(process.out).match() } ) } } }

关于该结构的几个关键点:

  • input[0]、input[1]、…按位置绑定 process 的输入通道;
  • tag用于为测试打标签,既可按主题分组(samtools),也支持精确到组件(samtools/sort),配合--tag参数可筛选运行;
  • setup { }块可以运行前置 process 来产出输入(见下文"测试模块"一节);
  • params { }块设置参数,config "..."行为测试加载指定配置;
  • nf-core 强制要求真实测试之外还要有一个 stub 测试——在test(...)块内添加options "-stub",即可驱动模块的stub:脚本来做管道冒烟验证。

关于 meta map(元数据映射)的约定也值得注意:nf-core 约定在输入/输出元组中携带[ id:'test', single_end:false ]这样的元数据 map,且 meta 永远是元组的第一个元素(见 developing.md 的 "The meta map convention")。测试中的input[0]直接遵循了这一约定。

测试模块(process)

测试一个模块的核心是when/then两段式结构:when块供应输入,then块对结果进行断言。

当被测模块需要另一个模块的输出作为输入时,使用setup块先运行前置模块:

test("sort then index") { setup { run("SAMTOOLS_SORT") { script "../../sort/main.nf" process { """ input[0] = [ [id:'test'], file(params.test_data + 'test.bam', checkIfExists:true) ] """ } } } when { process { """ input[0] = SAMTOOLS_SORT.out.bam """ } } then { assert process.success assert snapshot(process.out).match() } }

在这个例子中:

  1. setup块通过run("SAMTOOLS_SORT")执行排序模块,用script指向其main.nf,并为其绑定输入;
  2. when块把SAMTOOLS_SORT.out.bam(命名输出通道)直接作为被测 process 的input[0];
  3. then块断言成功并比对快照。

when块中也可以混用直接赋值与通道引用,例如某模块的输入部分来自手工构造的元组、部分来自上游通道。这种"先 setup 产出、再 when 消费"的模式正是 nf-test 对有依赖链的模块的标准测试手法。

断言(Assertions)

then块内的断言决定了测试的判定逻辑。当有多个检查项时,务必用assertAll(...)包裹,这样所有失败会一次性全部报告,而不是在第一个失败处中断。

常用断言句柄与辅助函数:

表达式检查内容
process.success/process.failed任务成功 / 失败
process.exitStatus == 0退出码
process.out.<emit>某个命名输出通道的内容
process.out.bam.get(0)第一个发射的元素
workflow.success、workflow.trace.tasks().size()工作流结果 / 任务数量
path(process.out.bam[0][1]).exists()某个文件是否存在
snapshot(...).match()与存储的快照比对
assertContainsInAnyOrder(ch, [...])通道包含给定元素(不要求顺序)

两个容易踩坑的要点:

  • 没有assertContainsInOrder。如果需要对文件内容做有序或子串检查,直接读取行内容再断言,例如:
    assert path(out[0][1]).readLines().any { it.contains('Done') } assert path(out[0][1]).readLines().last().contains('completed')
  • 插件(plugins)可以扩展领域相关的断言能力,例如nft-bam(BAM 文件)、nft-vcf(VCF 文件)、nft-utils(通用工具)。nf-core 会在nf-test.config中启用这些插件。

一个带文件内容断言与插件快照的完整示例:

then { assertAll( { assert process.success }, { assert path(process.out.bam[0][1]).exists() }, { assert snapshot( bam(process.out.bam[0][1]).getSamLinesMD5(), process.out.versions ).match() } ) }

这里bam(...).getSamLinesMD5()来自nft-bam插件,对 BAM 解析出的 SAM 行计算 MD5,将其与versions.yml一起纳入快照比对——既验证了二进制产物内容稳定,又验证了版本报告正确。

关于版本报告,developing.md 指出当前主流的versions.yml模式通过 HEREDOC 写入并作为path "versions.yml", emit: versions输出;nf-core modules create新生成的模块则改用 topic 通道 +eval()捕获版本。无论哪种模式,模块必须 emit 版本通道,且测试中应纳入快照,这正是上例快照中包含process.out.versions的原因。

快照测试(Snapshot Testing)

快照测试是 nf-test 最核心的回归保护机制:

snapshot(x).match() // 序列化 x 并与 .nf.test.snap 文件比对

首次运行会记录快照(生成.nf.test.snap);此后每次运行若输出发生变化,测试即失败。这实现了"一次编写,持续防回归"。

使用快照的三条黄金法则:

  1. 只快照稳定内容:文件 MD5/校验和、versions.yml、列表长度等;绝不快照绝对路径或时间戳(它们每次运行都会变,导致虚假失败);
  2. 有意的变更用命令重新生成快照:nf-test test --update-snapshot;
  3. 一个测试内可命名多个快照:.match("bam")、.match("versions")分别命名,便于多产物分别追踪。

快照文件(.nf.test.snap)需要提交进仓库。nf-core 要求每个模块/子工作流提交通过的快照,nf-core modules lint会检查其存在性——详见下文"nf-core 集成"。

测试工作流与流水线

对于整条流水线,使用nextflow_pipeline作用域,直接驱动main.nf:

nextflow_pipeline { name "Test full pipeline" script "../main.nf" test("default params") { when { params { outdir = "$outputDir"; input = "tests/samplesheet.csv" } } then { assert workflow.success assert workflow.trace.succeeded().size() > 0 } } }

关键差异与要点:

  • when块内通过params { }注入流水线参数,$outputDir是 nf-test 提供的自动输出目录变量;
  • then块断言workflow.success(整个流水线成功)以及workflow.trace.succeeded().size() > 0(确有任务成功完成);
  • nextflow_workflow作用域介于两者之间,用于测试子工作流(包含多个模块的复用单元),其写法与nextflow_pipeline相似,但script指向子工作流的main.nf。

测试数据要小:优先从 nf-core/test-datasets 选取微型输入,例如:

nf-core test-datasets search <term> # 在 nf-core/test-datasets 中查找小型测试文件

小数据意味着测试运行快、依赖少,可以频繁在 CI 中执行。这也呼应了 SKILL.md 中"Alwaystestfirst"的最佳实践——用-profile test,docker在真实数据之前先验证环境与管道。

运行测试:命令全览

nf-test 的 CLI 支持从"全量运行"到"单文件/单标签/变更集"的精细控制:

nf-test test # 运行全部测试 nf-test test modules/nf-core/samtools/sort/ # 运行某个目录下的测试 nf-test test tests/main.nf.test # 运行单个测试文件 nf-test test --tag samtools # 按标签筛选(配合文件内 tag 声明) nf-test test --profile docker # 选择容器引擎 nf-test test --update-snapshot # 接受新快照(重新生成 .nf.test.snap) nf-test test --changed-since HEAD^ # 只测试自某次提交以来变更的组件 nf-test test --only-changed --ci # CI 模式:缺失快照时失败而不是写入

各参数的作用:

  • --profile指定执行 profile(如docker、singularity、conda),与 Nextflow 的 profile 机制对接;
  • --update-snapshot用于接受因有意变更而失效的快照;
  • --changed-since <ref>只运行自该 ref 以来发生变更的组件,大幅缩短 CI 耗时;
  • --only-changed --ci是 CI 专用组合:如果快照缺失,测试失败而不是自动写入——防止 CI 静默生成未审查的快照。

nf-test.config在其中承担三项核心职责:设置测试目录、默认 profile、启用插件(nf-core 在此启用nft-bam/nft-vcf/nft-utils等)。

CI 的典型形态:只跑变更组件(--changed-since+--only-changed),配合 nf-core 官方提供的 nf-test GitHub Action 完成集成。这样每次 PR 只验证相关模块,快且准。

nf-core 集成:用包装命令管理测试

在 nf-core 生态中,优先使用nf-core工具的包装命令而不是直接调用nf-test,因为包装命令自动携带正确的 profiles、标签和快照处理逻辑:

nf-core modules create mytool # 脚手架:自动生成 tests/main.nf.test 等文件 nf-core modules test mytool # 运行该模块的 nf-test 套件 nf-core subworkflows test mysubwf

对应关系(详见 nf-core-tools.md):

命令作用
nf-core modules create [tool]脚手架新模块(main.nf、meta.yml、tests/)
nf-core modules test <tool>运行模块的 nf-test 套件
nf-core modules lint <tool>按模块规范 lint(校验测试文件与快照存在性)
nf-core subworkflows create/test/lint子工作流的同生命周期管理

nf-core 的硬性要求总结如下(来自 testing.md 与 developing.md):

  1. 每个 nf-core 模块/子工作流必须附带通过的 nf-test 测试;
  2. 测试必须包含一个stub 测试(用options "-stub"驱动stub:脚本);
  3. 快照(.nf.test.snap)必须提交进仓库;
  4. nf-core modules lint会检查这些测试产物的存在性,不满足则 lint 失败。

从 developing.md 的模块解剖可知,stub:块在模块中同样是被强制要求的:每个输出通道至少要touch出一个文件(gzip 输出则用echo '' | gzip > x.gz)。stub 测试因此能以近乎零成本的方式验证管道连接是否正确——这正是"stub 测试必须与真实测试并存"的设计初衷。

测试驱动的 nf-core 开发闭环

把以上内容串成一个完整的开发工作流(融合 nf-core-tools.md 的典型开发循环):

nf-core pipelines create # 脚手架流水线(自带 test profile 与 CI) nf-core modules install fastqc # 优先复用社区模块 nf-core modules create mytool # 新增自定义模块(自动生成测试脚手架) nf-core modules test mytool # 编写并运行 nf-test(真实 + stub 测试) nf-core test-datasets search <term> # 为测试挑选微型数据 nf-core modules lint mytool # 校验测试文件与快照齐备 nf-core pipelines lint # 流水线级 lint prettier --write . # 格式化(Harshil 对齐风格)

随后在 CI 中,通过--changed-since+--only-changed只验证变更组件,配合 nf-core 的 nf-test GitHub Action 完成自动回归。这样,从模块脚手架、测试编写、lint 校验到 CI 集成的整条链路都由工具保障——这也是 nf-test 成为 Nextflow 生态"唯一标准测试框架"的根本原因。

小结

nf-test 通过nextflow_process/nextflow_workflow/nextflow_pipeline三种作用域覆盖了 Nextflow 组件的全部测试层次:模块测试用when/then驱动、setup链式组装、assertAll聚合断言、快照机制防回归、-stub冒烟验证管道、--changed-since支撑高效 CI。结合 nf-core 工具链的create/test/lint生命周期与强制提交快照的规范,你可以把"测试"从一个事后动作变成组件开发的内置环节。

本文对应的完整技能入口位于 skills/nextflow/SKILL.md,更深入的内容可继续阅读同目录下的 language.md(DSL2 语言)、developing.md(nf-core 组件开发规范)与 nf-core-tools.md(nf-core CLI 完整参考)。

【免费下载链接】scientific-agent-skillsTurn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 190,000+ scientists worldwide. 165 ready-to-use validated skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.项目地址: https://gitcode.com/GitHub_Trending/cl/scientific-agent-skills

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询