评测与基准arXiv 2609.35922
VoxParity 评测:语音智能体在需要听声判断时仍按文字行事
提出 VoxParity 基准,评测语音智能体是否按音频改变工具调用
- arXiv
- 2609.35922
- 发表
- 层
- 应用层
- 场景
- AI Agent
- 测试
- MiMo-V2.6-Pro、gemini-3.7-flash、Qwen3.8-Omni 等 15 个
- 风险Agent 危险操作
- 组件音频
摘要VoxParity 用固定转写的音频对比和 words-only null,测量语音智能体是否按所听内容改变工具调用。
VoxParity 测量语音智能体在词语不变、只有声音改变时,是否按行业规则改它所执行的工具调用。