Conductor 向量数据库配置指南:多实例编排、参数详解与故障排查
【免费下载链接】conductorConductor is an event driven agentic workflow engine providing durable and highly resilient execution engine for applications and AI Agents项目地址: https://gitcode.com/GitHub_Trending/co/conductor
本指南系统讲解 Conductor(事件驱动的 Agentic 工作流引擎)中向量数据库(Vector Database)的配置体系。文章以 ai/VECTORDB_CONFIGURATION.md 为骨架,深入结合
ai模块源码,覆盖 PostgreSQL(pgvector)、MongoDB(Atlas Vector Search)、Pinecone 三种提供方的多命名实例配置格式、工作流引用方式、参数表、旧配置迁移与最佳实践。读完本文,你将能够为 RAG(检索增强生成)与语义搜索类工作流正确配置多套向量库实例,并能够独立排查常见故障。
概述:为什么需要可配置的向量数据库
Conductor 的 AI 能力(如LLM_STORE_EMBEDDINGS、LLM_SEARCH_EMBEDDINGS等系统任务)需要持久化与检索文本的向量表示(embeddings)。这些向量数据存储在独立的向量数据库中,而不同业务场景(文档索引、语义搜索、推荐召回)对环境、模型维度、延迟的要求各不相同。
为此,Conductor 采用多命名实例(multiple named instances)的配置方式,允许同时配置同一类型的多套数据库,以实现:
- 同类型多实例:例如同时连接多套 PostgreSQL 实例;
- 环境隔离:分别连接生产(prod)、开发(dev)、预发(staging);
- 用例分离:按用途拆分 embeddings 存储、搜索、推荐等不同向量库。
从源码来看,这一设计体现在 VectorDBInstanceConfig.java 中:配置类以@ConfigurationProperties(prefix = "conductor.vectordb")绑定外部配置,将每个命名实例按其type分派到对应的数据库实现,最终在 VectorDBProvider.java 中以ConcurrentHashMap<String, VectorDB>按名称注册,供工作流任务按名称查找。
支持的向量数据库
当前仓库支持以下三类向量数据库提供方:
| 类型标识(type) | 数据库 | 扩展/服务 |
|---|---|---|
postgres | PostgreSQL | pgvector 扩展 |
mongodb | MongoDB | Atlas Vector Search |
pinecone | Pinecone | 托管向量数据库服务 |
配置格式:conductor.vectordb.instances
所有向量数据库统一使用列表(list)式配置,位于conductor.vectordb.instances前缀之下。基础模板如下:
conductor: vectordb: instances: - name: "instance-name" # 该实例的唯一标识 type: "database-type" # 类型:postgres / mongodb / pinecone <type-specific-config>: # 与该数据库类型对应的配置块 # ... 类型专属属性每个列表元素对应一个独立实例,包含三个字段:
name:实例唯一名称,工作流任务通过它引用实例(要求与工作流中的vectorDB参数完全一致);type:数据库类型标识,决定实例化哪个实现类;- 类型专属配置块:按
type对应postgres、mongodb或pinecone键,提供该类型的具体参数。
在源码中,VectorDBInstanceConfig.VectorDBInstance 正是这样建模的:每个实例同时持有name、type以及三种类型的可选配置对象;createVectorDB 根据type(大小写不敏感)分派创建对应数据库实例。若类型未知或对应的配置块缺失,该实例会被记录错误日志并跳过。
配置示例
单个 PostgreSQL 实例
conductor: vectordb: instances: - name: "postgres-main" type: "postgres" postgres: datasourceURL: "jdbc:postgresql://localhost:5432/vectors" user: "conductor" password: "secret" dimensions: 1536 connectionPoolSize: 10 indexingMethod: "hnsw" # 可选:hnsw, ivfflat distanceMetric: "cosine" # 可选:l2, cosine, inner_product tablePrefix: "conductor"多个 PostgreSQL 实例
同一类型可配置多个实例,便于区分生产与开发环境(注意二者维度不同,需与各自使用的嵌入模型匹配):
conductor: vectordb: instances: - name: "postgres-prod" type: "postgres" postgres: datasourceURL: "jdbc:postgresql://prod-db:5432/vectors" user: "conductor" password: "prod-secret" dimensions: 1536 - name: "postgres-dev" type: "postgres" postgres: datasourceURL: "jdbc:postgresql://dev-db:5432/vectors" user: "conductor" password: "dev-secret" dimensions: 768MongoDB Atlas Vector Search
conductor: vectordb: instances: - name: "mongodb-embeddings" type: "mongodb" mongodb: connectionString: "mongodb+srv://user:pass@cluster.mongodb.net/" database: "conductor" collection: "embeddings" numCandidates: 100Pinecone
conductor: vectordb: instances: - name: "pinecone-search" type: "pinecone" pinecone: apiKey: "your-pinecone-api-key"混合配置(多种类型并存)
三种类型可以在同一个instances列表中任意混用,适合"生产检索 + 嵌入存储 + 缓存"的分工:
conductor: vectordb: instances: - name: "postgres-prod" type: "postgres" postgres: datasourceURL: "jdbc:postgresql://prod:5432/vectors" user: "conductor" password: "secret" dimensions: 1536 - name: "pinecone-embeddings" type: "pinecone" pinecone: apiKey: "pk-xxx" - name: "mongodb-cache" type: "mongodb" mongodb: connectionString: "mongodb://localhost:27017" database: "conductor"使用 properties 文件配置(等价写法)
除 YAML 外,也可以使用标准的application.properties索引式写法。仓库文档 docs/devguide/cookbook/ai-llm.md 给出了一个可直接用于 RAG 示例的配置:
conductor.vectordb.instances[0].name=postgres-prod conductor.vectordb.instances[0].type=postgres conductor.vectordb.instances[0].postgres.datasourceURL=jdbc:postgresql://localhost:5432/vectors conductor.vectordb.instances[0].postgres.user=conductor conductor.vectordb.instances[0].postgres.password=secret conductor.vectordb.instances[0].postgres.dimensions=1536两种写法最终都会被@ConfigurationProperties(prefix = "conductor.vectordb")绑定到同一套配置模型。
在工作流中使用向量数据库实例
工作流中的向量数据库系统任务通过inputParameters.vectorDB按配置的实例名称引用实例。以LLM_STORE_EMBEDDINGS为例(原文档 示例):
{ "name": "store_embeddings", "taskReferenceName": "store_embeddings_ref", "type": "LLM_STORE_EMBEDDINGS", "inputParameters": { "vectorDB": "postgres-prod", "index": "documents", "namespace": "my_namespace", "embeddings": "${embedding_task.output.embeddings}", "metadata": { "documentId": "${workflow.input.docId}" } } }从源码看,这一引用链路的实现位于 VectorDBWorkers.java:
LLM_STORE_EMBEDDINGS:把embeddings写入指定实例的index/namespace;LLM_SEARCH_EMBEDDINGS:基于查询向量检索相似文档,返回IndexedDoc列表,输入字段定义在 VectorDBInput.java(vectorDB、index、namespace、embeddings、query、metadata、maxResults等);- 两个任务最终都经由 VectorDBs.java 调用
VectorDBProvider.get(vectorDBName, context)按名称取出实例;若实例不存在,会抛出NonRetryableException("VectorDB not found: " + name),任务将直接失败而不会重试。
另外还有LLM_INDEX_TEXT(自动为文本生成嵌入并索引,见 VectorDBWorkers.java)以及LLM_SEARCH_INDEX、LLM_GET_EMBEDDINGS(已标记@Deprecated,内部转发到新方法)等任务。详细的 AI 工作流示例可参考 ai/examples 目录(如05-semantic-search.json、06-rag-basic.json、07-rag-complete.json)。
PostgreSQL 配置选项
| 属性 | 类型 | 默认值 | 说明 |
|---|---|---|---|
datasourceURL | String | 必填 | JDBC 连接 URL |
user | String | 必填 | 数据库用户名 |
password | String | 必填 | 数据库密码 |
dimensions | Integer | 256 | 向量维度 |
connectionPoolSize | Integer | 5 | 连接池大小 |
indexingMethod | String | "hnsw" | 索引方法(hnsw 或 ivfflat) |
distanceMetric | String | "l2" | 距离度量(l2、cosine、inner_product) |
invertedListCount | Integer | 100 | IVFFlat 索引参数 |
tablePrefix | String | null | 表名前缀 |
这些默认值在 PostgresConfig.java 中直接定义。结合 PostgresVectorDB.java 实现,有几个值得注意的底层行为:
- 连接池:
connectionPoolSize最终作用于 HikariCP(HikariDataSource)的maximumPoolSize,并设置idleTimeout为 60 秒;连接池启动时会等待其就绪(最长 5 秒,见 PostgresVectorDB.java)。 - 维度校验:写入时会校验
dimensions与传入向量的长度一致,不一致会直接抛出RuntimeException("Embeddings must be of dimensions: ...")。因此dimensions必须与嵌入模型输出的维度一致,否则运行时必然报错。 - 命名空间与表结构:
namespace与indexName必须匹配正则[a-zA-Z0-9_-]+,否则抛 "Invalid namespace/index name" 异常;表名由tablePrefix(若设置)与namespace拼接而成(tablePrefix + "_" + namespace),表不存在时自动创建。 - 索引与写入:写入前会自动创建向量表与向量索引,插入采用
ON CONFLICT (id) DO UPDATE的 upsert 语义;doc文本会经过TextUtils.sanitizeForPostgres清洗,metadata以 JSON 形式存储。 - 连接复用:DataSource 以
datasourceURL为键缓存在 Guava Cache 中,最多 100 个、空闲 60 秒后过期并关闭底层连接池(见 PostgresVectorDB.java)。
MongoDB 配置选项
| 属性 | 类型 | 默认值 | 说明 |
|---|---|---|---|
connectionString | String | 必填 | MongoDB 连接串 |
database | String | 必填 | 数据库名 |
collection | String | 可选 | 集合名 |
numCandidates | Integer | 可选 | 向量搜索参数(候选数) |
字段定义位于 MongoDBConfig.java,实际实现见 MongoVectorDB.java,其测试用例位于 MongoVectorDBTest.java。需要特别说明:向量搜索依赖 MongoDB Atlas 或 MongoDB 6.0+ 的 Atlas Search 能力,本地 MongoDB 容器不支持向量搜索,且必须先在集合上创建向量搜索索引。
Pinecone 配置选项
| 属性 | 类型 | 默认值 | 说明 |
|---|---|---|---|
apiKey | String | 必填 | Pinecone API Key |
字段定义位于 PineconeConfig.java,实现见 PineconeDB.java。使用前需确保 Pinecone 账号中已存在目标 index,且 API Key 具有相应权限。
从旧配置迁移
旧格式(每类型单实例)
在引入命名实例之前,配置采用单实例平铺结构:
conductor: vectordb: postgres: datasourceURL: "jdbc:postgresql://localhost:5432/vectors" user: "conductor" password: "secret"新格式(命名实例)
conductor: vectordb: instances: - name: "pgvectordb" # 使用旧类型名以保持向后兼容 type: "postgres" postgres: datasourceURL: "jdbc:postgresql://localhost:5432/vectors" user: "conductor" password: "secret"类型标识已简化:
pgvectordb→postgresmongovectordb→mongodbpineconedb→pinecone
不过,为了向后兼容,旧类型名仍然可用——只需将实例命名为对应旧类型名即可。这一点在源码中有直接印证:createVectorDB 的 switch 分支同时接受postgres/pgvectordb、mongodb/mongovectordb、pinecone/pineconedb两套写法(大小写不敏感)。
最佳实践
- 使用描述性名称:实例名应清晰表达用途(如
postgres-prod、pinecone-embeddings-search),便于在日志与工作流中识别。 - 隔离环境:生产、开发、预发使用不同实例,避免数据意外混写。
- 优化维度:
dimensions必须与嵌入模型输出维度一致,否则运行时抛错(PostgreSQL 实现会在写入时严格校验)。 - 连接池调优:根据工作负载与数据库容量调整
connectionPoolSize;过高会压垮数据库,过低会导致高并发下排队。 - 索引选型:
hnsw(默认)查询性能更好,适合在线检索;ivfflat建索引更快,适合写入密集、查询规模可控的场景,可通过invertedListCount调节其质量。 - 距离度量选型:
cosine适合归一化后的嵌入(绝大多数现代嵌入模型的默认场景);l2(欧氏距离)适合需要绝对距离语义的场景;inner_product用于点积相似度。所选度量应与训练/索引时一致。
故障排查
实例找不到(Instance Not Found)
出现Vector DB instance not found: xyz之类的错误时,依次检查:
- 工作流
inputParameters.vectorDB中的名称与配置中的name完全一致(区分大小写、无多余空格); - 实例确实已写入
application.yml/application.properties(注意conductor.vectordb.instances前缀是否正确); - 配置修改后是否已重启应用——
VectorDBProvider在启动阶段一次性完成实例注册(见 VectorDBProvider.java),运行期修改配置不会生效。
另外可留意服务启动日志:注册成功会打印Initialized vector DB instance: <name> (type: <type>),失败则打印Failed to initialize vector DB instance: ...及原因;VectorDBProvider启动时也会打印全部可用实例清单,方便核对名称。
PostgreSQL 连接问题
- 确保已安装 pgvector 扩展:
CREATE EXTENSION vector; - 核对 JDBC URL 格式与网络连通性(源码要求连接串非空,否则抛 "Missing connection URL" 异常);
- 检查数据库用户权限(建表、建索引、读写目标 schema 的权限)。
MongoDB 向量搜索问题
- 向量搜索需要 MongoDB Atlas 或 MongoDB 6.0+ 的 Atlas Search 能力;
- 确保已在集合上创建向量搜索索引(否则查询会失败);
- 本地 MongoDB 容器不支持向量搜索,请使用 Atlas 或兼容实例。
Pinecone 问题
- 验证 API Key 有效且具备必要权限;
- 确保目标 index 已存在于 Pinecone 账号中,再于工作流中引用。
小结
Conductor 的向量数据库配置围绕conductor.vectordb.instances这一统一列表展开,通过name + type + 类型专属配置块的结构同时支持 PostgreSQL、MongoDB 与 Pinecone 的多实例与混合部署;工作流侧以实例名解耦存储位置与业务逻辑,底层由VectorDBProvider统一注册、按名查找。理解本文中的参数语义、旧配置迁移路径与排查要点后,你即可为 RAG、语义搜索等 AI 工作流搭建稳定、可扩展的向量存储层。更多端到端示例可继续阅读 ai/examples 与 docs/devguide/cookbook/ai-llm.md。
【免费下载链接】conductorConductor is an event driven agentic workflow engine providing durable and highly resilient execution engine for applications and AI Agents项目地址: https://gitcode.com/GitHub_Trending/co/conductor
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考