☰
从一篇个人 Markdown 笔记到可检索语义索引:Khoj 文档接入处理链路实战解析
2026/10/11 4:03:54 网站建设 项目流程
  • 人工智能
  • AI 应用
  • 大模型
  • RAG
  • 后端
  • AI Agent

【免费下载链接】khoj

Your AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.

项目地址:https://gitcode.com/GitHub_Trending/kh/khoj
点击查看免费下载

导读

本文以 Khoj 仓库测试数据集中的一篇典型个人笔记(tests/data/markdown/Meet Arun and Pablo for Lunch.markdown)为贯穿线索,完整拆解 Khoj 将 Markdown 笔记转化为可被 AI 搜索与聊天引用的语义条目的处理链路:从文件解析、标题拆分、条目构造,到向量化入库与日期索引,再到用自然语言 +dt日期过滤器回查。读完本文,你将掌握 Khoj 的 Markdown 内容管线(MarkdownToEntries)的核心机制,并能在自托管部署中预判笔记如何被切分、索引与检索。

一、样本文件剖析:个人笔记中常见的三类信息载体

Khoj 定位为 "Your AI second brain",其核心能力之一就是把散落在本地磁盘、桌面同步目录中的个人 Markdown 文件变成可供语义搜索和聊天引用的知识库。仓库测试数据中收录了一批模拟真实个人笔记的文件,其中Meet Arun and Pablo for Lunch.markdown极具代表性,全文如下:

--- SCHEDULED: 2023-04-01 CLOSED: 2023-04-01 --- Met Pablo and Arun for Lunch at Arak, Medellin. Arun just sold his apartment in Nairobi and is moving with his wife to Medellin in April 2023! Pablo mentioned his son Amal just got admission into the Colegio Superior de Gastronomia in Mexico City. Last of his 3 kids to leave the nest! 2023-04-01 "Arak" "Dosa for Lunch" Expenses:Food:Dining 11.00 USD

这个文件浓缩了个人知识管理场景中三类高频信息载体:

  1. YAML/Org 风格 frontmatter 元数据:SCHEDULED/CLOSED字段记录了事件的计划与完成日期,是 Khoj 日期索引的天然素材;
  2. 自由文本正文:自然语言记录的见闻与事实,是语义检索的主要对象;
  3. 账本格式交易行(Ledger/Beancount 风格):2023-04-01 "Arak" "Dosa for Lunch"与Expenses:Food:Dining 11.00 USD,同样携带可被解析的结构化日期。

Khoj 会逐行处理这些内容,而不同类型的字段会走不同的子管线。下面依次展开。

二、入口与主流程:MarkdownToEntries 的处理骨架

在 Khoj 的源码中,Markdown 文件接入的统一入口是MarkdownToEntries.process(),位于 markdown_to_entries.py。该方法接收一个{文件路径: 文件内容}的字典,依次完成三个阶段:

def process(self, files: dict[str, str], user: KhojUser, regenerate: bool = False) -> Tuple[int, int]: deletion_file_names = set([file for file in files if files[file] == ""]) files_to_process = set(files) - deletion_file_names ... max_tokens = 256 with timer("Extract entries from specified Markdown files", logger): file_to_text_map, current_entries = MarkdownToEntries.extract_markdown_entries(files, max_tokens) with timer("Split entries by max token size supported by model", logger): current_entries = self.split_entries_by_max_tokens(current_entries, max_tokens) with timer("Identify new or updated entries", logger): num_new_embeddings, num_deleted_embeddings = self.update_embeddings(...)

三个阶段的职责分别是:

阶段方法作用
条目抽取extract_markdown_entries按标题结构把单个 Markdown 文件切成一个或多个Entry(raw 原文 + 行号 URI)
超长拆分split_entries_by_max_tokens对超过模型 token 上限的 compiled 文本做递归切块,保证每条不超过 256 tokens
增量入库update_embeddings计算 MD5 哈希、生成向量嵌入、增量写入数据库、索引日期

一个值得注意的细节:process中默认max_tokens = 256,而测试用例(test_markdown_to_entries.py)在直接调用extract_markdown_entries时常用max_tokens=3强制触发递归拆分,以便验证拆分边界行为。也就是说,256 是生产默认值,而 3/10/12 是测试专用的小值,目的是把"大文件拆小"的逻辑暴露出来。

2.1 空文件即删除信号

入口处还有一个易被忽略的约定:deletion_file_names = set([file for file in files if files[file] == ""])——内容为空的文件路径会被解释为"删除该文件的索引",不会进入抽取流程,而是最终在update_embeddings阶段通过EntryAdapters.delete_entry_by_file删除对应条目。这意味着客户端删除文件时,只需向服务端上报一个空内容条目即可触发索引清理。

三、条目切分原理:按标题层级递归、保留祖先标题

extract_markdown_entries遍历每个文件,调用process_single_markdown_file(markdown_to_entries.py)完成真正的拆分。其核心算法如下:

  1. 先拼接标题祖先:把当前 section 的所有祖先标题按层级顺序用#前缀拼在正文之前(ancestry_string),保证每个条目自带上文语境;
  2. 判定是否直接成条:若拼接后的内容 token 数 ≤max_tokens,且正文中不存在比当前层级更深的子标题,则整个内容作为一个条目;
  3. 否则递归拆分:按"下一个存在的标题层级"用正则re.split(rf"(\n|^)(?=[#]{{{next_heading_level}}} .+\n?)", ...)切出多个 section,对每个 section 更新其标题祖先(current_ancestry[next_heading_level] = current_section_title),再递归处理,同时用current_line_offset精确追踪每个 section 在文件中的起始行号。

对应的行为在测试中有非常明确的断言(test_markdown_to_entries.py):

def test_extract_entries_with_non_incremental_heading_levels(tmp_path): ... assert entries[1][0].raw == "# Heading 1\n#### Sub-Heading 1.1", "Ensure entry includes heading ancestory" assert entries[1][1].raw == "# Heading 1\n## Sub-Heading 1.2", "Ensure entry includes heading ancestory"

即:子标题条目会完整继承其所有祖先标题,即使标题层级跳跃(#直接跳到####)也能正确处理。

3.1 无标题文件的处理:恰适用于我们的样本文件

Meet Arun and Pablo for Lunch.markdown全文没有任何#标题,因此它的处理路径对应测试 test_extract_markdown_with_no_headings:

def test_extract_markdown_with_no_headings(tmp_path): ... # Ensure raw entry with no headings do not get heading prefix prepended assert not entries[1][0].raw.startswith("#") # Ensure compiled entry has filename prepended as top level heading assert entries[1][0].compiled.startswith(expected_heading)

可以推断,这个样本文件被抽取后会形成一个单一条目:raw保持原文不加标题前缀;而compiled(真正喂给嵌入模型和 LLM 的字段)会被自动加上文件名作为一级标题。这一点在convert_markdown_entries_to_maps中有直接实现(markdown_to_entries.py):

heading = parsed_entry.splitlines()[0] if re.search(r"^#+\s", parsed_entry) else "" # Append base filename to compiled entry for context to model prefix = f"# {entry_filename}\n#" if heading else f"# {entry_filename}\n" compiled_entry = f"{prefix}{parsed_entry}"

于是入库后的compiled大致形如:

# Meet Arun and Pablo for Lunch.markdown Met Pablo and Arun for Lunch at Arak, Medellin. ...

文件名作为顶级标题这一设计有实际价值:当 AI 引用该条目时,模型能直接看到内容来源文件名,回答中可自然带出"来自某篇笔记"的上下文。

3.2 行号溯源:每个条目携带精确文件定位

在 convert_markdown_entries_to_maps 中,每个条目还会生成形如file://{绝对路径}#line=N的 URI,其中N是条目在源文件中的起始行号(1 基)。本地文件走file://前缀,URL 则原样保存。

test_line_number_tracking_in_recursive_split 专门验证了这一点:它用max_tokens=10强制对大文件做多层递归拆分,然后逐个回读原始文件,校验每个条目 URI 中的行号所指向的行内容与条目首行一致。这意味着无论文件被拆得多碎,用户都能从检索结果跳回到笔记中的准确位置——这是"第二大脑"类产品体验的关键细节。

四、超长条目的二次拆分与清洗

当单个条目的 compiled 文本超过模型 token 上限时,split_entries_by_max_tokens(text_to_entries.py)会兜底处理。它使用langchain_text_splitters.RecursiveCharacterTextSplitter,切分优先级为:

text_splitter = RecursiveCharacterTextSplitter( chunk_size=max_tokens, separators=["\n\n", "\n", "!", "?", ".", " ", "\t", ""], keep_separator=True, length_function=lambda chunk: len(TextToEntries.tokenizer(chunk)), chunk_overlap=0, )

切块顺序是段落 > 行 > 感叹号 > 问号 > 句号 > 空格 > 制表符 > 字符,tokenizer即简单的text.split()(按空白切词)。额外的细节包括:

  • 从原始raw文本中反查每个 chunk 的实际位置,重算#line=行号,保证切块后的 URI 依然精确(text_to_entries.py);
  • 除首个 chunk 外,后续 chunk 会前置"截短的条目标题"(取标题最后 100 字符,text_to_entries.py),让模型知道这些碎片属于同一主题;
  • 用remove_long_words丢弃超过 500 字符的超长单词、用clean_field清除\0等非法字符,避免污染嵌入质量。

五、日期信息如何被提取与利用

样本文件中的SCHEDULED: 2023-04-01、CLOSED: 2023-04-01以及账本行首的2023-04-01都属于 Khoj 的日期索引体系。日期提取由DateFilter承担(date_filter.py),它维护了 20 种正则,覆盖结构化日期与自然语言日期:

  • 结构化:\b\d{4}[-\/]\d{2}[-\/]\d{2}\b(如2023-04-01)、\d{2}[-\/]\d{2}[-\/]\d{4}(如01-04-1984)等;
  • 自然语言:1st April 1984、April 2021、Apr 84等。

对应的行为有测试直接验证(test_date_filter.py):

extracted_dates = DateFilter().extract_dates("head CREATED: today SCHEDULED: 1984-04-01 tail") assert extracted_dates == [datetime(1984, 4, 1, 0, 0, 0)], "Expected only Y-m-d structured date to be extracted"

可见SCHEDULED: 2023-04-01这种写法会被稳定地解析出2023-04-01这个日期(相对日期如today会被忽略)。抽取到的日期在update_embeddings尾部被批量写入EntryDates表(text_to_entries.py),与条目建立关联,构成按时间检索的倒排索引。

5.1 查询侧:用 dt 过滤器按时间回查

入库的日期在查询侧通过dt过滤器消费。DateFilter定义了查询语法(date_filter.py):

dt([:><=]{1,2})"日期表达式"

常见用法包括:

查询示例语义
dt:"2023-04-01"命中 2023-04-01 当天
dt>="yesterday" dt<"tomorrow"昨天 00:00 至明天 00:00 之间
dt:"2 years ago"两年前的当天
dt:"last week"上一自然周

比较符<、<=、>、>=、=、:会被组合成交集区间(date_filter.py)。extract_date_range的测试给出了精确语义,例如dt:"1984-01-01"返回[当天0点, 次日0点)(test_date_filter.py)。

因此,对本文的样本笔记,若你自托管 Khoj 并已同步该文件,搜索:

dt:"2023-04-01" 午餐

即可精准命中"2023 年 4 月 1 日那次与 Arun、Pablo 的午餐"相关条目。更完整的过滤器说明可参考 query-filters.md 与 search.md。

六、增量更新与向量入库:哈希驱动的幂等同步

update_embeddings(text_to_entries.py)是管线的收尾环节,其增量策略非常清晰:

  1. 对每条compiled文本计算MD5 哈希(hash_func,见 text_to_entries.py);
  2. 从数据库按user + hashed_value + file_type查出已有哈希,只对新增哈希生成嵌入(embeddings_model[model.name].embed_documents(...));
  3. 将新条目批量写入Entry表,字段包括raw、compiled、heading(截断到 1000 字符)、file_path、hashed_value、corpus_id、url(即file://...#line=N)与search_model;
  4. 反过来删除数据库中已不存在于当前文件内容的旧哈希条目(to_delete_entry_hashes = existing - current);
  5. 同步更新FileObject中保存的文件原文,供后续全文检索使用。

这一设计意味着:重复同步同一文件是幂等的——内容未变则哈希不变、不产生新嵌入;内容局部修改则只有受影响条目被重算。process方法还会把新增/删除的条目数返回给调用方(markdown_to_entries.py),方便上层记录同步状态。

值得注意,MarkdownToEntries并非只服务本地文件:GitHub 数据源处理器也复用了它的process_single_markdown_file与convert_markdown_entries_to_maps(见 github_to_entries.py),说明按标题拆分与条目构造是 Khoj 所有 Markdown 类内容源的公共底座。

七、验证与测试:如何用仓库证据复现这条链路

如果你想亲手验证以上机制,仓库内已备好全部材料:

  • 样本输入:tests/data/markdown/ 目录下 20 个真实风格的笔记文件,除本文的主角外,还有带SCHEDULED/CLOSED的 Preparing to File Taxes for 2022.markdown、纯账本格式的 Miscellaneous Transactions.markdown、以及用于行号追踪验证的超长文件 main_readme.md;
  • 单元测试:test_markdown_to_entries.py 覆盖无标题、单条目、多条目、不同层级标题、非递增标题层级、标题前文本、小文件单条、递归拆分行号追踪共 8 类场景;
  • 日期测试:test_date_filter.py 覆盖SCHEDULED:解析、dt比较符区间、自然语言日期解析等;
  • 运行方式:仓库使用 pytest 组织测试(pytest.ini),可执行pytest tests/test_markdown_to_entries.py tests/test_date_filter.py运行上述用例。

八、小结:一条笔记的完整旅程

以Meet Arun and Pablo for Lunch.markdown为样本,Khoj 的完整处理链路可归纳为:

  1. 抽取:process_single_markdown_file判断文件无标题且体量小(< 256 tokens),整个文件作为单一条目,raw保留原文;
  2. 构造:convert_markdown_entries_to_maps以文件名为一级标题生成compiled,并生成file://...#line=1的定位 URI;
  3. 日期索引:DateFilter.extract_dates从SCHEDULED/CLOSED/账本行中抽出2023-04-01,写入EntryDates;
  4. 向量化:update_embeddings对 compiled 文本计算 MD5 与语义向量,增量写入Entry表;
  5. 检索:查询时用自然语言向量召回 +dt日期过滤器做时间限定,最终由 AI 引用对应条目作答。

理解了这条链路,你便能在自托管 Khoj 时预测笔记文件的索引行为:正文是否有标题决定条目切分粒度,frontmatter 中的结构化日期决定时间检索能力,文件是否超长决定是否二次切块。若希望让某类笔记获得更细粒度的检索(按小节命中),只需在笔记中合理使用##标题即可——这背后正是 markdown_to_entries.py 的标题递归拆分逻辑在起作用。

  • 人工智能
  • AI 应用
  • 大模型
  • RAG
  • 后端
  • AI Agent

【免费下载链接】khoj

Your AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.

项目地址:https://gitcode.com/GitHub_Trending/kh/khoj
点击查看免费下载

相关推荐

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询