Notableevaluation benchmarks

SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models

Kanpat Vesessook, Saksorn Ruangtanusak

Published
Sep 30, 2026 — 17:13 UTC

Problem

This work addresses the gap in evaluating spoken mathematical reasoning specifically in multi-turn speech-to-speech systems. The authors highlight the need for a structured benchmark to assess how well these systems can handle incremental spoken disclosures, which is critical for applications requiring conversational reasoning. The paper is a preprint and has not undergone peer review.

Method

The authors propose the SpeechConversationBench (SCB) framework, which utilizes a dataset of 103 sharded GSM8K problems. The evaluation conditions include three distinct setups: 1) Full - where the original problem is presented in a single turn; 2) Concat - where information shards are concatenated; and 3) Sharded - where the problem is disclosed incrementally across multiple turns. The evaluation involves four commercial speech systems alongside LEGO, a proprietary speech pipeline developed by the authors.

Results

The results indicate that the LEGO system achieved an accuracy of 77.5%, outperforming GPT-4o Realtime, which recorded an accuracy of 76.6%. Additionally, the study reports a significant accuracy decrease in the sharded condition, with reductions ranging from 5.0 to 25.3 percentage points compared to the concatenated condition. These results underscore the challenges faced by speech systems in maintaining accuracy across multi-turn interactions.

Limitations

The authors do not report any limitations in their study. However, the absence of reported limitations may suggest a need for further exploration of potential biases in the dataset or the generalizability of the results across different domains or languages.

Why it matters

The introduction of SCB provides a crucial tool for evaluating multi-turn reasoning capabilities in speech-to-speech models, which is essential for advancing conversational AI. By establishing a benchmark that focuses on incremental reasoning, this work lays the groundwork for future research aimed at improving the robustness and accuracy of speech systems in complex conversational scenarios.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI