☰
LlamaIndex 集成 TimescaleVectorStore:基于 PostgreSQL 的向量存储与时间感知相似检索实战指南
2026/10/10 18:55:31 网站建设 项目流程
  • 人工智能
  • RAG
  • 大模型

【免费下载链接】llama_index

LlamaIndex is the document processing platform for AI

项目地址:https://gitcode.com/GitHub_Trending/ll/llama_index
点击查看免费下载

本指南以 LlamaIndex 官方 API 参考文档 TimescaleVectorStore 为线索,围绕其完整源码实现与官方示例 Notebook,系统讲解如何用 TimescaleDB 作为 LlamaIndex 的向量存储后端。你将掌握TimescaleVectorStore的安装、参数配置、索引构建、ANN 索引管理,以及最具特色的"按时间分区的向量相似度检索"在 Retriever 与 Query Engine 中的落地用法。

TimescaleVectorStore 是什么

TimescaleVectorStore是 LlamaIndex 官方向量存储集成之一,其定位在官方集成清单中被列为TimeScale (TimescaleVectorStore),详见 vector_stores.md。它把 TimescaleDB——一个面向 AI 应用增强的 PostgreSQL——接入 LlamaIndex 的向量索引体系:

  • 增强 pgvector:通过受 DiskANN 启发的索引算法,在海量向量上提供更快、更准的相似度检索;
  • 自动按时间分区:向量及其元数据按时间自动分表,支持"既按向量相似度、又按时间范围"的高效检索;
  • 熟悉的 SQL 接口:向量与关系型元数据共存于同一数据库,可直接用 PostgreSQL 生态工具查询。

整个集成由独立包llama-index-vector-stores-timescalevector提供,核心实现只有约 300 行,位于 base.py,底层依赖timescale-vectorPython 客户端与llama-index-core,见 pyproject.toml。

安装与快速开始

安装依赖

在项目中安装 TimescaleVectorStore 集成包即可,示例 Notebook 中通常还会一并安装 OpenAI 嵌入模型包(见 Timescalevector.ipynb):

pip install llama-index-vector-stores-timescalevector pip install llama-index-embeddings-openai # 生成向量嵌入时使用

获取连接串并初始化

TimescaleVectorStore 的核心构造参数是service_url(Timescale 云数据库连接串)与table_name(存储向量的表名)。官方推荐通过from_params类方法创建实例,这是源码中定义的工厂方法(base.py):

from llama_index.vector_stores.timescalevector import TimescaleVectorStore TIMESCALE_SERVICE_URL = "postgres://tsdbadmin:<password>@<id>.tsdb.cloud.timescale.com:<port>/tsdb?sslmode=require" vector_store = TimescaleVectorStore.from_params( service_url=TIMESCALE_SERVICE_URL, table_name="your_table_name_here", num_dimensions=1536, # 可选,默认 1536 )

也可以在环境变量中存放连接串(.env中以TIMESCALE_SERVICE_URL=postgresql://开头),再用python-dotenv读取,避免密钥硬编码。

参数说明

from_params与构造函数__init__的参数完全一致(base.py):

参数类型默认值说明
service_urlstr必填PostgreSQL/Timescale 连接串
table_namestr必填存储向量的表名,构造时会被自动转为小写
num_dimensionsintDEFAULT_EMBEDDING_DIM(1536)向量维度,须与所用嵌入模型输出维度一致
time_partition_intervalOptional[timedelta]None时间分区间隔;传入后启用按时间分区能力,且表 id 必须是 UUID v1

初始化时会自动完成两件事(base.py):

  1. _create_clients():分别创建同步客户端client.Sync与异步客户端client.Async,两者共享同一连接串、表名与维度。若设置了time_partition_interval,id 类型固定为UUID,否则为TEXT;
  2. _create_tables():通过同步客户端调用create_tables()自动建表,无需手工执行 DDL。

与 VectorStoreIndex 集成:建索引、查询、复用

从文档构建索引

TimescaleVectorStore 可像其他向量存储一样作为VectorStoreIndex的后端。通过StorageContext把向量存储注入索引,LlamaIndex 会负责切分文档、调用嵌入模型并写入 Timescale 表:

from llama_index.core import SimpleDirectoryReader, StorageContext, VectorStoreIndex from llama_index.vector_stores.timescalevector import TimescaleVectorStore # 加载文档 documents = SimpleDirectoryReader("./data/paul_graham").load_data() # 创建向量存储 vector_store = TimescaleVectorStore.from_params( service_url=TIMESCALE_SERVICE_URL, table_name="paul_graham_essay", ) # 构建索引 storage_context = StorageContext.from_defaults(vector_store=vector_store) index = VectorStoreIndex.from_documents(documents, storage_context=storage_context) # 查询 query_engine = index.as_query_engine() response = query_engine.query("Did the author work at YC?")

复用已有的索引

Timescale 表中的数据持久保存在云端,因此重启进程后只需连接串与表名即可恢复索引,无需重新入库(Timescalevector.ipynb 中的 "Querying existing index" 一节):

vector_store = TimescaleVectorStore.from_params( service_url=TIMESCALE_SERVICE_URL, table_name="paul_graham_essay", ) index = VectorStoreIndex.from_vector_store(vector_store=vector_store) query_engine = index.as_query_engine() response = query_engine.query("What did the author do before YC?")

源码级实现剖析

数据写入:node 转行

add/async_add负责把 LlamaIndex 的BaseNode转换成 Timescale 客户端可识别的行并upsert入库(base.py):

  • 元数据经node_to_metadata_dict序列化,remove_text=True表示文本单独存列、不重复嵌入元数据;flat_metadata决定元数据是否扁平化存储;
  • 默认复用node.node_id作为主键;
  • 启用时间分区后:主键必须是 UUID v1。源码会先尝试解析现有 id,若非 UUID 或不是 v1 版本,则用uuid.uuid1()自动生成,从而让行的"时间属性"由插入时刻决定;
  • 每一行包含[id, metadata, text, embedding]四个字段。

查询与元数据过滤

query/aquery将VectorStoreQuery转换为底层搜索(base.py):查询向量取query.query_embedding,返回条数取query.similarity_top_k。VectorStoreQuery是 LlamaIndex 核心定义的通用查询结构,含query_embedding、similarity_top_k、filters等字段,见 types.py。

MetadataFilters会被_filter_to_dict拍平成{key: value}字典传给底层搜索;空过滤器返回None,不做额外条件。结果通过_db_rows_to_query_result还原为VectorStoreQueryResult(nodes / similarities / ids),优先用metadata_dict_to_node重建节点,若旧数据格式不兼容则回退到TextNode兼容逻辑,保证向后兼容。

删除

delete(ref_doc_id)以{"doc_id": ref_doc_id}作为元数据过滤条件调用delete_by_metadata,即按文档 ID 批量清理其下所有节点向量(base.py)。

同步/异步双通道

存储类同时维护client.Sync与client.Async两套底层客户端,同步方法(add/query/delete)与异步方法(async_add/aquery)一一对应,方便在异步应用中使用;close()会同时关闭两个客户端(base.py)。

加速检索:三种 ANN 索引的管理

当数据量增长后,可以对 embedding 列创建 ANN 索引加速相似度检索。注意这里的"索引"是数据库层的 ANN 索引,与 LlamaIndex 的索引概念不同。TimescaleVectorStore 通过枚举IndexType支持三种索引(base.py):

枚举值底层索引特点
TIMESCALE_VECTOR(默认)DiskANN 启发的图索引TimescaleVectorIndex默认推荐,海量向量下检索更快更准
PGVECTOR_HNSWpgvector 的 HNSW(分层可导航小世界图)召回精度高,官方同样推荐
PGVECTOR_IVFFLATpgvector 的 IVFFLAT(倒排文件索引)适合大数据量批量导入后使用

创建与删除

create_index(index_type=DEFAULT_INDEX_TYPE, **kwargs)默认创建 timescale_vector(DiskANN)索引;drop_index()删除当前索引(base.py):

# 默认创建 DiskANN 索引 vector_store.create_index() # 删除后,用自定义参数重建 DiskANN 索引 vector_store.drop_index() vector_store.create_index("tsv", max_alpha=1.0, num_neighbors=50) # 切换为 HNSW 索引(m、ef_construction 有智能默认值) vector_store.drop_index() vector_store.create_index("hnsw", m=16, ef_construction=64) # 切换为 IVFFLAT 索引(num_lists、num_records 有智能默认值) vector_store.drop_index() vector_store.create_index("ivfflat", num_lists=20, num_records=1000)

注意事项

  • 单表单索引:PostgreSQL 中一张表的一个列只能有一个索引。若想对比不同索引类型的性能,可以建多张表、在同一张表加多个向量列分别建索引,或反复 drop 后重建对比;
  • 时机建议:最好在数据大部分入库之后再创建 ANN 索引(如 IVFFLAT 依赖数据分布统计);
  • 推荐取舍:官方示例建议日常优先使用timescale-vector(DiskANN)或HNSW索引。

时间感知检索:TimescaleVectorStore 的核心差异化能力

这是 Timescale Vector 区别于普通 pgvector 方案的关键特性:向量与元数据按时间自动分区,检索时可以同时约束"向量相似度 + 时间范围",并且只扫描相关分区,效率很高。典型应用场景包括:LLM 对话历史存储与召回、按最近时间检索相似新闻、对知识库做时间段限定问答。

1. 启用时间分区

在创建 store 时传入time_partition_interval,例如按 7 天一个分区:

from datetime import timedelta ts_vector_store = TimescaleVectorStore.from_params( service_url=TIMESCALE_SERVICE_URL, table_name="li_commit_history", time_partition_interval=timedelta(days=7), )

分区粒度需按查询习惯权衡:频繁查最近数据可用timedelta(days=1),跨十年的大时间窗可用半年或一年。

2. 让节点携带历史时间戳

启用分区后主键必须是 UUID v1。如果节点代表过去某时刻的数据(例如 git 提交记录),需用时间戳手工生成 UUID v1;如果希望绑定"当前时间",则无需处理——入库时源码会自动生成 UUID v1。示例 Notebook 中通过timescale_vector客户端的uuid_from_time完成:

from timescale_vector import client from datetime import datetime def create_uuid(date_string: str): time_format = "%a %b %d %H:%M:%S %Y %z" datetime_obj = datetime.strptime(date_string, time_format) return str(client.uuid_from_time(datetime_obj)) # 构造带历史时间戳 id 的 TextNode node = TextNode( id_=create_uuid(record["date"]), text=record_content, metadata={"commit": ..., "author": ..., "date": ...}, )

3. 三种时间过滤查询方式

query()额外接受时间过滤关键字参数,内部经date_to_range_filter构造client.UUIDTimeRange(base.py):

关键字含义
start_date起始时间(含边界由start_inclusive控制,默认含)
end_date结束时间(含边界由end_inclusive控制,默认含)
time_delta时间增量,配合start_date或end_date使用
start_inclusive/end_inclusive边界是否包含,默认均为包含

方法一:给定起止日期范围:

from llama_index.core.vector_stores import VectorStoreQuery vector_store_query = VectorStoreQuery(query_embedding=query_embedding, similarity_top_k=5) query_result = ts_vector_store.query( vector_store_query, start_date=start_dt, end_date=end_dt )

方法二:从起始日期向后推一个时间窗(start_date + time_delta):

query_result = ts_vector_store.query( vector_store_query, start_date=start_dt, time_delta=timedelta(days=7) )

方法三:从结束日期向前推一个时间窗(end_date - time_delta):

query_result = ts_vector_store.query( vector_store_query, end_date=end_dt, time_delta=timedelta(days=7) )

三种方式返回的都是VectorStoreQueryResult,其中nodes只包含落在指定时间范围内的向量,且只扫描相关分区,查询效率很高。

4. 在 Retriever 与 Query Engine 中使用时间过滤

时间过滤参数可以通过vector_store_kwargs透传给 Retriever 与 Query Engine,从而把时间段限定直接融入 RAG 链路(Timescalevector.ipynb 第四节):

from llama_index.core import VectorStoreIndex index = VectorStoreIndex.from_vector_store(ts_vector_store) # Retriever 限定时间窗 retriever = index.as_retriever( vector_store_kwargs={"start_date": start_dt, "time_delta": timedelta(days=7)} ) nodes = retriever.retrieve("What's new with TimescaleDB functions?") # Query Engine 限定时间窗 query_engine = index.as_query_engine( vector_store_kwargs={"start_date": start_dt, "end_date": end_dt} ) response = query_engine.query("What's new with TimescaleDB functions?")

这样即可实现"只基于某个时间段内的知识回答近期问题"等时间感知 RAG 应用。

使用流程总结

  1. 在 Timescale 云平台创建 PostgreSQL 数据库,获取形如postgres://tsdbadmin:<password>@<id>.tsdb.cloud.timescale.com:<port>/tsdb?sslmode=require的service_url;
  2. pip install llama-index-vector-stores-timescalevector(如需嵌入,一并安装对应 embedding 包);
  3. 通过from_params(service_url=..., table_name=..., num_dimensions=..., time_partition_interval=...)创建实例,表会自动创建;
  4. 用StorageContext将 store 注入VectorStoreIndex,从文档建索引,或通过from_vector_store复用已有表;
  5. 数据量上来后用create_index创建 DiskANN / HNSW / IVFFLAT 索引加速检索;
  6. 需要时间感知检索时,设置time_partition_interval、使用 UUID v1 节点 id,并在query、as_retriever、as_query_engine中传入start_date/end_date/time_delta。

延伸阅读

  • 集成包完整源码:base.py
  • 官方实战示例:Timescalevector.ipynb
  • 通用向量存储协议与查询结构:types.py
  • 向量存储集成总览:vector_stores.md
  • 人工智能
  • RAG
  • 大模型

【免费下载链接】llama_index

LlamaIndex is the document processing platform for AI

项目地址:https://gitcode.com/GitHub_Trending/ll/llama_index
点击查看免费下载

相关推荐

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询