☰
如何为qwen-audio-agent接入新语音模型:自定义Realtime Provider开发完全教程
2026/10/2 6:54:27 网站建设 项目流程

如何为qwen-audio-agent接入新语音模型:自定义Realtime Provider开发完全教程

【免费下载链接】qwen-audio-agentA realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents项目地址: https://gitcode.com/gh_mirrors/qw/qwen-audio-agent

qwen-audio-agent 是一个为 AI Agent 打造的实时语音运行时,让你的 Agent 持续“开口、干活、在场”。内置前台覆盖了 DashScope、GPT-Live、Google Live、豆包 Seeduplex 等主流实时语音模型,而本文将以新手视角,带你快速掌握 qwen-audio-agent 的自定义 Realtime Provider 开发流程:不改动通用会话与后台逻辑,只需写一个独立适配器,就能把任意新语音模型接入语音 Agent。

1️⃣ 先搞懂:Realtime Provider 在项目中的位置

qwen-audio-agent 采用分层架构:前台语音服务负责“说话”,Gateway 负责编排会话、工具调用与任务,后台 Agent 负责“干活”。Realtime Provider 正是前台与 Gateway 之间的适配器:

  • 每个 Provider 是独立适配器,完整拥有自己的 URL、认证、模型、Session 和错误分类语义;
  • Gateway 的工具调用、任务、客户端协议完全不变,你只需翻译事件格式;
  • 官方扩展指南:docs/voice-frontends/custom-provider.zh.md。

内置 Provider 一览(位于 server/src/voice/providers/):

Provider key服务源码文件
dashscope通义 Qwen 实时语音/全模态dashscope.mjs
gpt-liveGPT-Live(GA 协议)gpt-live.mjs
google-liveGoogle Live APIgoogle-live.mjs
doubao-seeduplex豆包 Seeduplex 全双工doubao-seeduplex.mjs
minicpm-o/stepfun/s2sMiniCPM-o / 阶跃 / Speech-to-Speechproviders 目录

2️⃣ 最快上手方法:Provider 契约清单

注册时会严格校验,缺少任一成员都会立即抛错。Provider 对象必须提供:

成员作用
key/label唯一标识(小写字母数字与连字符)与展示名
inputSampleRate/outputSampleRate输入/输出 PCM 采样率(数值)
url()/headers()服务地址与认证头,可从宿主配置闭包读取令牌
model()/voice()当前模型与音色
isConfigured()是否已配置(决定 Provider 是否出现在列表中)
classifyError()将错误信息归类为inactivity、fatal、response_slot_busy等
buildSession()生成session.update载荷(指令、工具、音频格式、VAD)
buildSpeakResponse()主动播报请求的载荷
buildResultInjection()/buildPermissionInjection()把任务结果、权限请求注入为对话项 + 响应指令
protocol或createProtocol()事件协议适配器(每条连接可生成独立实例)

协议适配器(protocol)需要实现的核心方法(校验逻辑见 provider-registry.mjs):

  • 出站:encodeOutgoing、sessionUpdate、audioAppend、imageAppend、conversationItemCreate、responseCreate、responseCancel
  • 入站:normalizeIncoming(把上游事件翻译成 Gateway 标准事件)
  • 关联:conversationItemId、correlateResponseCreate、responseCorrelationId
  • 构造辅助:userTextItem、functionOutputItem

💡 好消息:如果你的新模型兼容 OpenAI Realtime 协议族,可以直接复用现成的 GA 方言适配器 ga-protocol.mjs——它已处理output_modalities字段改名、文本增量事件映射、按类型命名空间化的 item ID 等细节,GPT-Live 就是靠它接入的。

3️⃣ 最小示例:一个自定义 Provider 长什么样

以 GPT-Live 为模板(gpt-live.mjs)精简后的核心形态:

export const myProvider = { key: 'my-vendor', label: 'My Vendor', aliases: ['mv'], // 可选:兼容旧写法 inputSampleRate: 24000, outputSampleRate: 24000, protocol: gaRealtimeProtocol, // 复用 OpenAI Realtime GA 协议 capabilities: { perResponseInstructions: true, // 按需声明,键必须在校验白名单内 }, model: () => config.myVendorModel, voice: () => config.myVendorVoice || null, isConfigured: () => Boolean(config.myVendorApiKey), url: () => `${config.myVendorUrl}?model=${config.myVendorModel}`, headers: () => ({ Authorization: `Bearer ${config.myVendorApiKey}` }), classifyError: message => /unauthorized|401/i.test(message) ? 'fatal' : 'other', buildSession: ({ agentContext, sessionOptions }) => ({ type: 'realtime', instructions: buildFrontendInstructions(agentContext), tools: myVendorTools(agentContext), output_modalities: ['audio'], audio: { input: { format: { type: 'audio/pcm', rate: 24000 }, turn_detection: {...} }, output: { format: { type: 'audio/pcm', rate: 24000 }, voice: sessionOptions?.voice }, }, }), buildSpeakResponse: content => ({ conversation: 'none', modalities: ['audio'], tool_choice: 'none', instructions: speakResponseInstructions(content), }), buildResultInjection: (content, { allowTools = false } = {}) => ({ item: gaRealtimeProtocol.userTextItem(content), response: { modalities: ['audio'], tool_choice: allowTools ? 'auto' : 'none' }, }), buildPermissionInjection: permission => ({ item: gaRealtimeProtocol.userTextItem(`<permission_request>...`), response: { modalities: ['audio'], tool_choice: 'none' }, }), }

协议差异(字段改名、事件映射、ID 命名空间)应尽量收敛在 protocol 里,而不是散落在业务代码中——这正是官方推荐的做法。

4️⃣ 注册到 Gateway:一步装配

通过createRealtimeProviderRegistry注入即可,无需改任何会话运行时代码:

import { createGatewayApplication } from 'qwen-audio-agent/gateway-application' import { createRealtimeProviderRegistry } from 'qwen-audio-agent/realtime-provider' import { myProvider } from './my-vendor-provider.mjs' const realtimeProviderRegistry = createRealtimeProviderRegistry({ providers: [myProvider], defaultProvider: 'my-vendor', }) createGatewayApplication({ realtimeProviderRegistry, realtimeProvider: 'my-vendor', })

几个实用选项:

  • visibility: 'gateway-only':Provider 仅供宿主选择,不出现在桌面设置页与公共列表;
  • validateSessionOptions:上游连接前的同步校验,只拒绝已确认无效的配置,未知值交给服务端验证;
  • connectionMessages():WebSocket 打开后、session.update之前发送原始握手帧。

5️⃣ 能力标志:正确声明服务的“怪癖”

Provider 通过capabilities布尔标志告知 Gateway 上游服务的真实行为,避免前台按名字分支判断。常用项(完整清单见 realtime-provider.mjs):

标志何时设为 false / true
acknowledgesConversationItems服务不返回 conversation-item 确认事件时设 false
conversationItemIdEcho服务端会重分配 item ID 时设 false,网关按唯一待确认项关联
restoreConversationContext注入历史会被当作实时用户输入时设 false
automaticToolResponses工具结果由服务原生续答时设 true(如豆包 Seeduplex、Google Live)
imageRequiresAudioStart首帧视频前必须先有音频时设 true,网关用 20ms 静音初始化时间线

6️⃣ 行为验证:跑一遍共享契约测试

接入新 Provider 后,在仓库根目录运行:

node --test server/test/realtime-provider-behavior.test.mjs

该测试(realtime-provider-behavior.test.mjs)通过真实会话运行时 + 本地 WebSocket 服务,覆盖全部 Provider 的上下文投递、工具续答、授权、取消、异步播报排队与重连恢复。新增协议时补充对应 fixture 再跑共享测试——这是可重复的契约验证,但不能替代真实服务联调与音频设备测试。

7️⃣ (可选)接入桌面设置页

仓库内置前台的设置统一定义在 shared/realtime-provider-definitions.mjs(不依赖 Node.js、不含密钥),有可选模型时同步维护 shared/realtime-model-catalog.mjs。

界面固定显示“服务地址、API Key、模型、音色”四行,分别对应endpoint、credential、model、voice配置槽位;桌面选择器、表单、配置读写都消费这份定义,无需新增供应商 HTML 面板。

⚠️ 注意:宿主注入的自定义 Provider 不会自动注册到桌面设置页,设置定义与运行时适配器是分离的。

✅ 小结:接入新语音模型的完整清单

  1. 读懂官方扩展文档,确认目标服务协议是否兼容 OpenAI Realtime 协议族;
  2. 实现 Provider 契约(10 个方法 + 2 个采样率),差异收敛进 protocol;
  3. 按真实行为声明capabilities标志;
  4. 用createRealtimeProviderRegistry装配进 Gateway;
  5. 跑realtime-provider-behavior共享测试 + 真实服务联调。

参考文件速查:

  • 扩展指南:docs/voice-frontends/custom-provider.zh.md
  • Provider 注册与校验:server/src/voice/providers/provider-registry.mjs、server/src/voice/providers/registry.mjs
  • 语音运行时入口:server/src/voice/
  • 行为验证测试:server/test/realtime-provider-behavior.test.mjs

照着这份清单走,你的新语音模型就能以“原生前台”的身份加入 qwen-audio-agent,与 WebUI、桌面客户端、移动端共享同一套会话能力。🎙️

【免费下载链接】qwen-audio-agentA realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents项目地址: https://gitcode.com/gh_mirrors/qw/qwen-audio-agent

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询