SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models
Kanpat Vesessook, Saksorn Ruangtanusak
- Published
- Sep 30, 2026 — 17:13 UTC
Problem
This work addresses the gap in evaluating spoken mathematical reasoning specifically in multi-turn speech-to-speech systems. The authors highlight the need for a structured benchmark to assess how well these systems can handle incremental spoken disclosures, which is critical for applications requiring conversational reasoning. The paper is a preprint and has not undergone peer review.
Method
The authors propose the SpeechConversationBench (SCB) framework, which utilizes a dataset of 103 sharded GSM8K problems. The evaluation conditions include three distinct setups: 1) Full - where the original problem is presented in a single turn; 2) Concat - where information shards are concatenated; and 3) Sharded - where the problem is disclosed incrementally across multiple turns. The evaluation involves four commercial speech systems alongside LEGO, a proprietary speech pipeline developed by the authors.
Results
The results indicate that the LEGO system achieved an accuracy of 77.5%, outperforming GPT-4o Realtime, which recorded an accuracy of 76.6%. Additionally, the study reports a significant accuracy decrease in the sharded condition, with reductions ranging from 5.0 to 25.3 percentage points compared to the concatenated condition. These results underscore the challenges faced by speech systems in maintaining accuracy across multi-turn interactions.
Limitations
The authors do not report any limitations in their study. However, the absence of reported limitations may suggest a need for further exploration of potential biases in the dataset or the generalizability of the results across different domains or languages.
Why it matters
The introduction of SCB provides a crucial tool for evaluating multi-turn reasoning capabilities in speech-to-speech models, which is essential for advancing conversational AI. By establishing a benchmark that focuses on incremental reasoning, this work lays the groundwork for future research aimed at improving the robustness and accuracy of speech systems in complex conversational scenarios.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
