WorldSonus: Bringing Sound to Worlds
Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang, Haoyuan Huang, Yaguang Wu, Xiangjun Huang, Ziyang Ma, Weijia Chen, Hongyu Liu, Zeyue Tian, Qifeng Chen
- Published
- Oct 6, 2026 — 17:50 UTC
Problem
This work addresses the challenge of real-time sound generation for interactive video streams, focusing on the need for interactive control of sound instructions during streaming. It emphasizes the importance of spatially aligned stereo sound that accurately reflects scene geometry and camera motion. The paper is a preprint and has not yet undergone peer review.
Method
The authors propose a streaming causal autoregressive diffusion architecture designed for real-time sound generation, achieving a real-time factor of 0.41. The interactive control mechanism is facilitated through an audio-centric captioning pipeline that employs chunk-indexed prompt scheduling. The model is trained on high-quality stereo supervision derived from a diverse set of stereo and ambisonic data, ensuring robust performance across various audio scenarios.
Results
The acoustic quality of the proposed system matches or outperforms state-of-the-art bidirectional models, although specific baseline comparisons are not provided. Similarly, the spatial alignment of the generated sound also matches or outperforms existing bidirectional models, with no explicit baseline mentioned.
Limitations
The authors do not report any limitations in their work, and no obvious limitations are identified in the available text.
Why it matters
The implications of this research are significant for the development of interactive media applications, where real-time sound generation can enhance user experience by providing immersive audio that is contextually relevant to the visual content. This work could pave the way for more advanced audio-visual integration techniques in gaming, virtual reality, and other interactive environments.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
