☰
OpenHuman 的 Tauri v2 iOS 插件 tauri-plugin-ptt 实战:推按即讲录音识别与 TTS 语音合成
2026/10/11 9:39:46 网站建设 项目流程

OpenHuman 的 Tauri v2 iOS 插件 tauri-plugin-ptt 实战:推按即讲录音识别与 TTS 语音合成

【免费下载链接】openhumanOpenHuman is an open source personal AI for Mac, Windows and Linux — local-first memory, agent orchestration, and deep research.项目地址: https://gitcode.com/GitHub_Trending/op/openhuman

OpenHuman 是一个面向 Mac、Windows 与 Linux 的开源个人 AI 应用,其移动端语音交互依赖一个独立的 Tauri v2 插件tauri-plugin-ptt实现"推按即讲(Push-to-talk)+ 语音合成(TTS)"能力。本文以该插件在仓库中的官方文档为主体,结合 Rust 命令层、Swift 原生实现与 JS 绑定源码,系统讲解它的命令/事件契约、iOS 权限配置、底层音频会话管理原理以及真机测试要点,帮助读者在自研 Tauri v2 移动应用中原样落地这套"按住说话、松开上屏、语音朗读"的完整链路。

插件定位:为 Tauri v2 补齐 iOS 语音能力

tauri-plugin-ptt是仓库内位于 packages/tauri-plugin-ptt 的独立 Tauri v2 插件,目标平台明确为iOS。它包装了三套系统框架:

  • AVAudioEngine—— 麦克风音频采集引擎;
  • Speech.framework(SFSpeechRecognizer)—— 语音转文字(STT);
  • AVSpeechSynthesizer—— 文字转语音(TTS)。

插件对外暴露 5 个命令、5 类异步事件和一套错误码,前端通过plugin:ptt|<command>形式调用。插件对桌面端(Mac/Windows/Linux)做了隔离降级:所有命令在非 iOS 目标上统一返回NotSupported错误,因此桌面构建完全不受影响(见 src/lib.rs 中PttHandle的 stub 分支)。

整体架构:JS → Rust → Swift 三层调用链

README 给出了一张清晰的架构图,结合源码可还原完整链路:

JS (MascotScreen) ↓ invoke / listen Rust (commands.rs) ↓ PluginHandle::run_mobile_plugin Swift (PTTPlugin.swift) ↓ PTTRecorder — AVAudioEngine + SFSpeechRecognizer PTTSpeaker — AVSpeechSynthesizer AudioSessionManager — AVAudioSession lifecycle + notifications

三层各司其职:

  1. JS 层:guest-js/index.ts 通过invoke('plugin:ptt|<name>')调用命令,通过listen('ptt://<event>')订阅事件;
  2. Rust 层:src/commands.rs 定义 5 个#[command],经tauri::generate_handler!注册;src/mobile.rs 用tauri::ios_plugin_binding!生成 Swift↔Rust 的 FFI 胶水,每个命令经PluginHandle::run_mobile_plugin将载荷序列化为 JSON,调用PTTPlugin上对应的@objc func;
  3. Swift 层:ios/Sources/tauri-plugin-ptt/PTTPlugin.swift 是 Tauri 插件类,聚合三个组件:PTTRecorder(录音+识别)、PTTSpeaker(合成)、AudioSessionManager(会话生命周期与系统通知)。

值得注意的是 Rust 命令命名(snake_case)到 Swift 方法(camelCase)的自动映射:start_listening→startListening,cancel_speech→cancelSpeech等,Swift 侧注释明确要求命令名必须与 commands.rs 一致。

命令契约(Commands)

Command描述
start_listening激活AVAudioEngine+SFSpeechRecognizer。部分识别结果以事件流形式持续到达。
stop_listening停用录音会话并返回最终识别文本。
speak入队一条AVSpeechSynthesizer语音合成任务。
cancel_speech立即停止当前合成任务。
list_voices列出所有AVSpeechSynthesisVoice.speechVoices()语音。

前端调用 API

JS 绑定位于 guest-js/index.ts,对应关系如下:

// 开始录音(部分转写通过事件异步到达) await startListening(); // 停止录音,返回最终文本 { text, isFinal } const result: TranscriptEvent = await stopListening(); // 合成语音:voiceId 可选,rate 为 0.5–2.0 的倍率(默认 1.0 = 正常语速) await speak('Hello from OpenHuman', { voiceId: 'com.apple.voice.compact.en-US.Samantha', rate: 1.2, }); // 立即停止合成 await cancelSpeech(); // 枚举设备语音 [{ id, name, lang }] const voices: VoiceInfo[] = await listVoices();

命令参数在 Rust 侧定义于 src/models.rs:SpeakRequest携带text、voice_id、rate三个字段,其中rate的取值范围注释为0.5(慢)~ 2.0(快),默认1.0;TranscriptResult返回text与恒为true的is_final(仅在stop_listening返回时成立)。

命令的 Swift 侧实现要点

在 PTTPlugin.swift 中:

  • startListening异步执行recorder.startListening(),成功则invoke.resolve(),失败则invoke.reject(...)并额外向事件总线发出ptt://error(权限拒绝或音频错误);
  • stopListening同步取出最终文本,先触发ptt://transcript-final事件,再返回TranscriptResult(text:isFinal: true);
  • speak用invoke.parseArgs(SpeakArgs.self)解析参数后交给PTTSpeaker.speak;
  • listVoices将 Swift 字典映射为VoiceInfoPayload数组返回。

事件契约(Events)

所有事件经 Tauri 事件总线发往 "main" 目标,前端用listen订阅:

EventPayload描述
ptt://transcript-partial{ text: string }录音过程中实时返回的部分转写结果
ptt://transcript-final{ text: string }stop_listening之后的最终结果
ptt://tts-started{ utteranceId: string }合成开始
ptt://tts-ended{ utteranceId: string; finished: boolean }合成结束(finished: false表示被取消)
ptt://error{ code: string; message: string }异步错误(权限、中断等)

JS 侧提供了 5 个订阅函数,均返回UnlistenFn以便在组件卸载时取消订阅:

const unlistenPartial = await onTranscriptPartial(text => { // 实时更新按住说话期间的转写预览 }); const unlistenFinal = await onTranscriptFinal(text => { // 松手后把最终文本发送到会话 }); const unlistenStarted = await onTtsStarted(id => { /* 开始朗读 */ }); const unlistenEnded = await onTtsEnded((id, finished) => { // finished === false 表示被 cancelSpeech 打断 }); const unlistenErr = await onError(err => { // err: { code, message } });

事件在 Swift 侧如何触发

PTTPlugin.load(webview:)中把闭包回调接线到trigger:

  • recorder.onPartialTranscript→ 触发ptt://transcript-partial;
  • recorder.onError→ 触发ptt://error;
  • speaker.onStarted/speaker.onEnded→ 分别触发ptt://tts-started/ptt://tts-ended。

PTTSpeaker通过AVSpeechSynthesizerDelegate回调上报生命周期:didStart上报(uid, true),didFinish上报(uid, true),didCancel上报(uid, false)。utteranceId 由UUID().uuidString生成并记录在currentUtteranceId,由于插件同一时刻只入队一条合成任务,委托回调读取最近一次设置的 id 即可保证对应关系正确(见 PTTSpeaker.swift)。

错误码(Error codes)

Code触发场景
permission_denied麦克风或语音识别权限被拒绝
interrupted电话或系统音频打断了录音会话
route_changed录音过程中蓝牙耳机断开
audio_errorAVAudioEngine失败
recognition_errorSFSpeechRecognizer转写失败

错误码的产生路径在 Swift 侧分两条:

  1. 同步返回:Rust 侧Error枚举(见 src/error.rs)包含MicrophonePermissionDenied、SpeechPermissionDenied、AlreadyRecording、NotRecording、AudioEngine、SpeechRecognizer、Tts等变体,通过自定义Serialize实现把错误字符串化后返回给 JS;
  2. 异步事件:PTTPlugin.emitPermissionOrAudioError把PTTRecorder.RecorderError.microphonePermissionDenied/speechPermissionDenied统一映射为permission_denied,其余映射为audio_error并携带localizedDescription。

此外PTTRecorder的识别任务回调里做了专门的错误过滤:kAFAssistantErrorDomain下 code 209(用户取消)与 1110(无语音输入)被视为正常结束,不触发recognition_error,只有其他错误才上报(见 PTTRecorder.swift)。

iOS 权限配置(Info.plist)

首次调用startListening时系统会弹出权限对话框,因此必须在 Info.plist 中声明两个用途描述:

<key>NSMicrophoneUsageDescription</key> <string>Used for push-to-talk voice messages.</string> <key>NSSpeechRecognitionUsageDescription</key> <string>Used to transcribe your voice to text.</string>

权限请求逻辑在PTTRecorder.requestPermissions()中,同时兼容新旧 iOS API:

  • iOS 17+ 使用AVAudioApplication.requestRecordPermission;
  • 更早版本回退到AVAudioSession.sharedInstance().requestRecordPermission;
  • 语音识别统一通过SFSpeechRecognizer.requestAuthorization,状态必须为.authorized才继续。

录音识别与语音合成的底层实现

PTTRecorder:单会话式的录音 + STT 管线

PTTRecorder.swift 遵循"一次startListening创建一个识别任务,stopListening完整拆除"的单会话模型,任务绝不跨会话残留。关键细节:

  • 识别请求shouldReportPartialResults = true,实时回传部分转写;requiresOnDeviceRecognition = false,即允许走网络识别;
  • 通过engine.inputNode.installTap(onBus:0, bufferSize:1024, ...)把音频缓冲持续append到SFSpeechAudioBufferRecognitionRequest,全程不落盘;
  • stopListening()先调用request.endAudio()与task.finish()让识别器基于已缓冲内容完成收尾,再摘除 tap、停止引擎、解激活音频会话,并返回镜像的latestTranscript;
  • forceStop()用于应用退到后台或会话被中断的强停场景:task.cancel()后直接清理,不等待最终结果;
  • 识别任务回调会持续更新latestTranscript,供停止时读取——因为SFSpeechRecognitionTask本身不暴露result属性。

PTTSpeaker:语速映射与取消语义

PTTSpeaker.swift 对AVSpeechSynthesizer做薄封装:

  • 未指定voiceId时,使用设备当前语言的语音(AVSpeechSynthesisVoice(language: Locale.current...));
  • 语速映射:调用方传入的归一化倍率0.5–2.0(JS 侧1.0 = 正常)会被先clamp到[0.1, 2.0],再按AVSpeechUtteranceDefaultSpeechRate * clamped换算到 AVFoundation 的[0,1]语速刻度(AVFoundation 默认速率0.5对应调用方的1.0);
  • cancel()使用stopSpeaking(at: .immediate),随后委托回调didCancel上报finished: false,前端据此区分"读完"与"被打断"。

AudioSessionManager:录音/播放共享单一会话

AudioSessionManager.swift 以单例形式集中管理AVAudioSession,让录音与播放共享一套 category 配置,避免蓝牙场景下反复切换 category 引发爆音:

  • activateForRecording()设置category: .playAndRecord、mode: .spokenAudio,options 包含.defaultToSpeaker、.allowBluetooth、.allowBluetoothA2DP—— 这正是 README 测试清单中"iPhone 离开耳朵时默认走扬声器外放"的实现基础;
  • deactivate()以.notifyOthersOnDeactivation释放会话并通知其他应用恢复音频;
  • startObserving注册AVAudioSession.interruptionNotification(电话/系统音频抢占)与routeChangeNotification(蓝牙设备插拔)两个系统通知,回调交给PTTPlugin统一处理。

中断与路由变更的优雅降级

PTTPlugin对两类系统扰动做了兜底(见 PTTPlugin.swift 的handleInterruption/handleRouteChange/appDidBackground):

  • 电话中断:先stopListening()取得最终文本并触发ptt://transcript-final,再发出ptt://error(code: interrupted);
  • 蓝牙断开:仅在.oldDeviceUnavailable且正在录音时,同样先收尾再发route_changed错误;
  • 退到后台:应用didEnterBackground时若正在录音则停止并产出最终转写,同时speaker.cancel()释放音频会话。

权限系统:Tauri 能力(capabilities/permissions)

作为 Tauri v2 插件,命令调用受权限系统管控。permissions/autogenerated/reference.md 为 5 个命令各生成了一对allow/deny权限标识,例如:

  • ptt:allow-start-listening/ptt:deny-start-listening
  • ptt:allow-stop-listening/ptt:deny-stop-listening
  • ptt:allow-speak/ptt:deny-speak
  • ptt:allow-cancel-speech/ptt:deny-cancel-speech
  • ptt:allow-list-voices/ptt:deny-list-voices

在移动端 Tauri 应用的 capability 文件(仓库中见 src-tauri-mobile/capabilities)里为对应窗口授予所需权限即可,未授权的命令调用会被拒绝。

桌面端降级:no-op stub

插件在非 iOS 平台上不引入任何原生依赖:src/lib.rs 中的PttHandle<R>通过#[cfg(target_os = "ios")]条件编译——iOS 上持有PttMobile<R>,其他平台退化为PhantomData<fn(R) -> R>占位。占位类型特意选用函数指针fn(R) -> R而非PhantomData<R>,以保证结构体在R不满足Send + Sync时依然满足 Taurimanage()的Send + Sync + 'static约束。5 个方法在非 iOS 分支统一返回Error::NotSupported并打印 warn 日志,因此桌面构建可以安全引用该插件而不影响主流程。

手动测试清单(Manual testing checklist)

Swift 原生层无法在 CI 中做单元测试(需要 iOS 工具链与模拟器),因此官方文档要求按下列清单在真机或模拟器上逐项验收:

  • 首次调用startListening时弹出权限对话框;
  • 说话过程中部分转写实时更新,停止后最终转写与内容一致;
  • 按住按钮录音、松开停止,聊天消息携带转写文本发出;
  • iPhone 离开耳朵时,TTS 默认通过扬声器外放;
  • 蓝牙耳机音频路由正确;录音中断开耳机能优雅停止;
  • 录音过程中应用退到后台,能产出最终转写并干净停止;
  • 电话打断时发出ptt://error,code为interrupted;
  • TTS 播放中调用cancelSpeech,收到tts-ended且finished: false;
  • listVoices返回非空的AVSpeechSynthesisVoice列表。

其中第 4 条与第 5 条分别由上文介绍的.defaultToSpeaker会话选项与routeChangeNotification监听保证;第 8 条由didCancel委托回调保证。

工程结构与构建配置

插件按 Tauri v2 移动插件标准布局组织:

  • src:Rust 侧lib.rs/commands.rs/mobile.rs/models.rs/error.rs;
  • guest-js:JS 绑定源码与 Vitest 单测(index.test.ts 用 mock 的invoke/listen验证每个函数调用了正确的命令名、参数结构与事件名);
  • ios/Sources/tauri-plugin-ptt:PTTPlugin.swift/PTTRecorder.swift/PTTSpeaker.swift/AudioSessionManager.swift;
  • permissions:权限标识、schema.json与自动生成的reference.md。

构建配置方面:Cargo.toml 声明crate-type = ["cdylib", "rlib"]并依赖tauri 2、serde、thiserror,iOS 目标无需额外 Rust 依赖(桥接走ios_plugin_binding!路径);Package.swift 要求 iOS 16+ 并静态链接Tauri框架;package.json 以tauri-plugin-ptt-api作为 JS 包名,peer 依赖@tauri-apps/api >= 2.0.0。

小结

tauri-plugin-ptt给出了一个在 Tauri v2 iOS 应用中落地"推按即讲"与"语音合成"的完整范式:Rust 命令层负责跨端统一契约并在桌面端优雅降级,Swift 层以AVAudioEngine + SFSpeechRecognizer与AVSpeechSynthesizer承接系统能力,AudioSessionManager统一管理会话与系统通知,最后通过 5 类事件把实时转写、合成生命周期与异步错误流回传给前端。阅读 packages/tauri-plugin-ptt/README.md 可快速掌握契约全貌,深入对应源码则可复用到自研移动插件的权限申请、中断降级与事件桥接设计中。

【免费下载链接】openhumanOpenHuman is an open source personal AI for Mac, Windows and Linux — local-first memory, agent orchestration, and deep research.项目地址: https://gitcode.com/GitHub_Trending/op/openhuman

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询