☰
avoid-ai-writing 三级词表体系全解析:112 个 AI 高频词条如何分级、替换与豁免
2026/10/4 10:44:57 网站建设 项目流程

avoid-ai-writing 三级词表体系全解析:112 个 AI 高频词条如何分级、替换与豁免

【免费下载链接】agentsMulti-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity项目地址: https://gitcode.com/GitHub_Trending/agents24/agents

导读

本文深入剖析avoid-ai-writing插件中负责词汇审计的核心参考文档——word-tiers.md。该文档维护了 112 个单词/短语词条(分三个层级)外加 10 条多词样板短语,定义了"哪些词是 AI 写作信号、在什么强度下该替换、哪些场景下可以豁免"。读完本文,你将掌握这套分级词表的完整内容、每条词目的推荐替换方案、形态变体匹配规则、load-bearing等特殊词条的连字符陷阱与构造豁免,以及它在 detect/rewrite/edit 三种模式下与语境(context)配置协同工作的底层逻辑。


一、词表总览:层级表达的是"权重",不是"修改难度"

word-tiers.md开篇即点明这套词表的设计哲学:层级(tier)表示一次命中携带多少权重(weight),而不是修复难度。所有词条都可以被替换,且替换动作本身难度相同;真正不同的是"一个命中意味着什么"。

文档统计口径为:

  • 112 个单词与短短语词条,分布在三个层级;
  • 10 条多词样板短语(Tier 3 phrases);
  • 计数源自上游词表(upstream tables),并会在上游的 CI 中对这些数量强制校验("Counts here are derived from the upstream tables and enforced against them in the upstream CI")。

按插件 README.md 的说明,这套词表派生于 MIT 许可的conorbronsdon/avoid-ai-writing上游技能,其词汇研究基础来自brandonwise/humanizer——后者的研究"减少了对单独出现完全正常的词、却在聚集成簇时可疑的词的误报"(the tiering is adapted from brandonwise/humanizer's vocabulary research, which reduces false positives on words that are fine alone and suspicious in clusters)。

形态变体匹配规则

每个词条覆盖所列单词及其形态学变体:副词(adverb)、动名词(gerund)、复数(plural)、比较级(comparative)以及动词变位(verb conjugations)。例如:

  • delve覆盖delving;
  • leverage覆盖leveraging。

文档特别给出了一条语境判断原则:当某个变体承载了独立的、正当的含义时,应依据语境判断而非盲目匹配。典型示例是real——当它表示"真实的、事实性的"(如 "a real improvement" 表示真实的改进)时,它不再作为强化词(intensifier)处理。


二、Tier 1:始终替换(always replace)

Tier 1 又细分为两个"带"(band):1A与1B。两者都会被替换,且编辑动作完全相同;区别在于一次命中的含义。

2.1 Tier 1A:AI 频率标记词(49 条)

1A 收录的是"被声称在机器文本中出现的频率远高于人类写作"的词汇。成簇出现时,这些词构成关于一段文字产生方式的证据。49 条完整词表如下:

替换(Replace)建议写法(With)
delve / delve intoexplore, dig into, look at
landscape(隐喻用法)field, space, industry, world
tapestry(describe the actual complexity) 描述真实的复杂性
realmarea, field, domain
paradigmmodel, approach, framework
embarkstart, begin
beacon(rewrite entirely) 整体重写
testament toshows, proves, demonstrates
robuststrong, reliable, solid
comprehensivethorough, complete, full
cutting-edgelatest, newest, advanced
leverage(动词)use
pivotalimportant, key, critical
underscoreshighlights, shows
meticulous / meticulouslycareful, detailed, precise
seamless / seamlesslysmooth, easy, without friction
game-changer / game-changingdescribe what specifically changed and why it matters 描述具体改变了什么及其意义
hit differently / hits different(say what specifically changed, or cut) 说出具体变化,或删除
watershed momentturning point, shift(或描述发生了什么变化)
marking a pivotal moment(state what happened) 陈述实际发生的事
the future looks bright(cut — say something specific or nothing) 删除,改为说具体的事
only time will tell(cut — say something specific or nothing) 同上
nestledis located, sits, is in
vibrant(describe what makes it active, or cut)
thrivinggrowing, active(或引用一个数字)
despite challenges… continues to thrive(name the challenge and the response, or cut) 点名挑战与应对,或删除
showcasingshowing, demonstrating(或删除该从句)
deep dive / dive intolook at, examine, explore
unpack / unpackingexplain, break down, walk through
bustlingbusy, active(或引用使其繁忙的具体内容)
intricate / intricaciescomplex, detailed(或点名具体的复杂度)
complexities(name the actual complexities, or use "problems" / "details")
ever-evolvingchanging, growing(或描述如何变化)
enduringlasting, long-running(或引用持续了多久)
dauntinghard, difficult, challenging
holistic / holisticallycomplete, full, whole(或描述包含什么)
actionablepractical, useful, concrete
impactfuleffective, significant(或描述具体影响)
learningslessons, findings, takeaways
thought leader / thought leadershipexpert, authority(或描述其实际贡献)
best practiceswhat works, proven methods, standard approach
at its core(cut — just state the thing) 删除,直接陈述
synergy / synergies(describe the actual combined effect) 描述实际协同效果
interplayrelationship, connection, interaction
keen(作强化词时)interested, eager, enthusiastic(或删除,直接陈述兴趣)
genuinely / genuine(作强化词时)(cut — just state the fact) 删除,直接陈述事实
symphony(隐喻用法)(describe the actual coordination or combination)
embrace(隐喻用法)adopt, accept, use, switch to
load-bearing(隐喻用法)essential, critical, necessary —— 或说明移除它会导致什么失效
load-bearing的特殊规则:连字符是分水岭

文档为load-bearing单独标注了两条规则:

  1. 必须连字符:不带连字符的 "load bearing" 是普通英语("the load bearing down on the bridge"),只有连字符复合形式才是 AI 信号。
  2. 构造豁免(construction carve-out):当load-bearing出现在字面结构名词(wall,beam,column,joist,truss,footing,slab,stud,masonry,lintel,rafter,girder,capacity)之前——中间最多夹一个材质或位置形容词——它属于标准建筑术语,不标记。而抽象名词(structure,element,frame,foundation)仍保持可标记状态,因此 "the load-bearing structure of his argument" 依然会被命中。
关于 1A 的证据边界(Caveat)

文档明确提醒一个值得保留的注意事项:"在 AI 文本中出现频率远高于人类"这一说法是继承而来,而非实测所得。它追溯至brandonwise/humanizer,后者声明了一个 5 到 20 倍的比率,但没有公开其方法与数据集。因此在有人用机器写作语料实测这个比率之前,1A 应被视为"有充分支持的惯例(convention)",而不是经过测量验证的事实。

2.2 Tier 1B:清晰度编辑(10 条)

1B 处理的是冗长与虚浮的形式化表达(wordiness and inflated formality)。替换它们本身就是好写作,与作者是谁无关——一次 1B 命中不是机器作者身份的证据。文档给出的实测背景是:对照 257 段 2023 年之前经核实的人类散文,1B 词条在普通专业写作中的出现率相当可观——in order to、utilize、commence、ascertain、endeavor就是一部分人惯用的词汇。

因此 1B 的处理策略是:

  • 按 Tier 2 的权重处理(两个以上同段出现才标记),并将其排除在任何密度信号之外——这样一次冗余修复永远不会把一篇文档推向 AI 分类;
  • 在 detect 模式下,两个带必须分开报告(1A 是作者身份相关证据,1B 只是风格建议)。

10 条完整词表:

替换(Replace)建议写法(With)
utilizeuse
in order toto
due to the fact thatbecause
serves asis
features(动词)has, includes
boastshas
presents(虚浮用法)is, shows, gives
commencestart, begin
ascertainfind out, determine, learn
endeavoreffort, attempt, try

三、Tier 2:同一段落出现两个及以上才标记(40 条)

Tier 2 收录的词汇单独出现完全正常。两个或两个以上挤在同一段落里,通常意味着该段落需要重写。40 条完整词表:

替换(Replace)建议写法(With)
harnessuse, take advantage of
navigate / navigatingwork through, handle, deal with
fosterencourage, support, build
elevateimprove, raise, strengthen
unleashrelease, enable, unlock
streamlinesimplify, speed up
empowerenable, let, allow
bolstersupport, strengthen, back up
spearheadlead, drive, run
resonate / resonates withconnect with, appeal to, matter to
revolutionizechange, transform, reshape(或描述改变了什么)
facilitate / facilitatesenable, help, allow, run
underpinsupport, form the basis of
nuancedspecific, subtle, detailed(或点名具体的微妙之处)
crucialimportant, key, necessary
multifaceted(describe the actual facets, or cut)
ecosystem(隐喻用法)system, community, network, market
myriadmany, numerous(或给出数字)
plethoramany, a lot of(或给出数字)
encompassinclude, cover, span
catalyzestart, trigger, accelerate
reimaginerethink, redesign, rebuild
galvanizemotivate, rally, push
augmentadd to, expand, supplement
cultivatebuild, develop, grow
illuminateclarify, explain, show
elucidateexplain, clarify, spell out
juxtaposecompare, contrast, set side by side
paradigm-shifting(describe what actually shifted) 描述实际转变了什么
transformative / transformation(describe what changed and how)
cornerstonefoundation, basis, key part
paramountmost important, top priority
poised (to)ready, set, about to
burgeoninggrowing, emerging(或引用数字)
nascentnew, early-stage, emerging
quintessentialtypical, classic, defining
overarchingmain, central, broad
quietlycut,或点名具体的对比
deeply(仅限意义搭配:"deeply integrated"、"deeply committed"、"deeply rooted";字面用法如 "deeply nested" 或 "cares deeply" 永远不计入聚集)cut,或点名具体深入之处
underpinning / underpinningsbasis, foundation, what supports

注意deeply条目附带的限定:只有"意义类搭配"才算数,字面用法("deeply nested"、感情上的 "cares deeply")永远不参与聚集计数——这条边界防止了把正常的英文误判为 AI 信号。


四、Tier 3:仅在高密度时标记(13 条)

Tier 3 收录的是完全正常的词汇,仅在文本"饱和"时才标记。其逻辑是:当这些笼统赞美词填满了本该放置具体细节的空间时,说明文字出了问题。文档给出的参考阈值是约占总词数的 3%时值得再看一眼。

词汇处理方式
significant / significantly用具体内容替换部分用法:数字、对比、示例
innovative / innovation描述实际的新颖之处
effective / effectively说明方式或引用指标
dynamic / dynamics点名实际的驱动力或变化
scalable / scalability描述什么可扩展、扩展到什么程度
compelling说明它为什么有说服力
unprecedented点名它打破了什么先例(或删除)
exceptional / exceptionally引用使它成为例外的事实
remarkable / remarkably说明值得关注之处
sophisticated描述这种复杂性本身
instrumental说明它扮演的角色
world-class / state-of-the-art / best-in-class引用基准或对比
verbatim通常与动词冗余("copies X verbatim" = "copies X")——删除;若"逐字"标记形成对比,则点名它:byte-for-byte, word for word, unchanged。该词在法务/研究/QA 语境中是术语("verbatim transcript / record / testimony"),在该语境下先衡量密度再决定是否标记

五、Tier 3 短语:密度或聚集时标记(10 条)

多词样板短语(multi-word boilerplate)单独出现时毫无问题,但在生成内容中会大量堆叠。文档特别点名:Crypto、web3、DePIN 和 AI 基础设施类评测文章是重灾区。

针对这类短语有两条触发规则:

  1. 同一短语在一篇中使用了两次;
  2. 一篇中出现了来自本表的三个及以上不同短语(即使每个只出现一次)。

第二条规则专门捕捉模型的典型形态——模型通过"变化自己的样板话"来显得不那么重复,而这三个不同短语的聚集恰恰暴露了这一点。

短语处理方式
emerging sector / emerging space / emerging category点名实际的行业或它"新兴"在哪里
the integration of (X with Y)描述整合了什么、对用户有何改变
the intersection of (X and Y)选择真正重要的具体重叠点,或删掉这个框架
community-driven点名社区实际做了什么。"Community-driven" 单独出现是空话
long-term sustainability引用时间跨度与约束。"Long-term" 是空泛之词
user engagement点名具体行为。"Engagement" 是点击/评论/留存之上的包装
decentralized compute说明具体架构,或删除。该短语已成类别标签而非论断
(sustainable) reward emissions引用排放时间表与去向(sink)
tokenized incentive structures描述实际机制(vesting、gauge、bonded LP 等)
designed for long-term [X]删掉 "designed for"——要么是、要么不是。然后陈述该属性

六、语境调整:technical-blog / docs / casual 的豁免矩阵

词表不是一刀切的。word-tiers.md末尾给出了三层语境调整,完整的逐规则容差矩阵在 profiles.md 中:

6.1technical-blog:技术含义回归

在技术博客语境下,以下词汇恢复其正当技术含义,不再标记:

robust、comprehensive、seamless、ecosystem、leverage(当主语是真实的平台杠杆或 API 时)、facilitate、underpin、streamline

但即使在这种语境下,以下词仍然标记:

delve、tapestry、beacon、embark、testament to、game-changer、harness

6.2docs与casual

  • docs(文档、README、指南):放宽整张词表(relaxes the full table),清晰度优先于语气;
  • casual(聊天、内部笔记、快速回复):只降级到 P0 模式(drops to P0 patterns only),即只处理最严重的可信度杀手。

6.3 词表在技能中的实际调用位置

词表不是孤立存在的。按 SKILL.md 定义的审计流程(The pass),词汇检查是五步中的第三步:

  1. 选择语境配置(linkedin / technical-blog / investor-email / docs / casual / blog 默认);
  2. 扫描 pattern-catalog.md 中的 P0/P1 模式(快速检查覆盖 P0+P1,完整审计再加 P2);
  3. 对照 word-tiers.md 的分级词表检查词汇——Tier 1 默认替换(先应用语境配置的例外);Tier 2 在同一段落出现两个及以上时替换;Tier 3 仅在文本饱和时替换;
  4. 节奏检查放最后,权重最高——结构规律性在词汇替换后依然存活,统一句长、统一段长、对称句式比任何单个标记词都重要;
  5. 重写后再读一遍自己的重写——循环过渡语、系动词回避(copula avoidance)和新的虚浮词总是能在第一遍中幸存。

此外,profiles.md 的容差矩阵中,词表一栏(Word table)在各语境下的强度为:linkedin 严格、blog 严格、technical-blog部分(即上文 6.1 的豁免清单)、investor-email 严格、docs 放宽、casual 仅 P0。Tier 3 短语聚集(Tier 3 phrase clustering)在 investor-email 中为额外严格——一封投资邮件里出现一个 "thriving ecosystem" 就可能削弱整封邮件的可信度。


七、使用边界:词表是写作质量信号,不是作者身份判决

贯穿整个插件的立场(详见 SKILL.md 开篇)是:这些模式在模型输出中更常见,但人类——尤其是赶稿、在不熟悉的体裁中写作或使用第二语言的人——也会大量使用。机器检测的证据是双向的:一份 Stanford 审计发现七个检测器将 61% 的非母语英语写作者的 TOEFL 作文标记为 AI 生成(而母语写作者约 5%);2025 年的一份审计发现开源检测器误报率在 30%–78% 之间,不适合高风险场景。

因此:

  • 每个 flag 都应被视为写作质量信号;
  • 该技能不对任何文本做作者身份分类,任何 flag 都不应决定学术诚信、招聘或署名问题;
  • 1A 的 "5–20x 比率" 是继承自上游的惯例而非实测,使用时需保留这一证据边界;
  • 在 detect 模式下,清晰度编辑(1B)必须与作者身份标记(1A)视觉分离——一次冗余修复说明不了谁写了这段文字,它只能作为风格建议。

结语:从词表到可执行的写作纪律

word-tiers.md提供的不是一份"禁用词黑名单",而是一套按权重分层的证据体系:1A 是频率标记、1B 是清晰度、Tier 2 是聚集信号、Tier 3 是密度信号、Tier 3 短语捕捉样板的自我伪装。配合 SKILL.md 的模式目录与 profiles.md 的语境矩阵,这套词表可以无缝嵌入 detect-only、rewrite、edit-in-place 三种模式,在 README、changelog、PR 描述、博客与社交文案等各类场景中落地。记住文档反复强调的那句原则:一个在语境中恰好正确的词,即使在词表里,也应当保留。

【免费下载链接】agentsMulti-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity项目地址: https://gitcode.com/GitHub_Trending/agents24/agents

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询