MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo
- Published
- Sep 30, 2026 — 16:31 UTC
Problem
The paper addresses a gap in the computational efficiency of content encoders in one-step voice conversion models. Existing models, such as MeanVoiceFlow, struggle with speed and quality, necessitating improvements for practical applications. This work is a preprint and has not undergone peer review.
Method
The authors introduce MeanVoiceFlow2, which employs a training method called conversion distillation. This method utilizes MeanVoiceFlow alongside real data reconstruction to enhance the model's performance. Additionally, the training incorporates techniques such as Diffusion-GAN training, which involves sample mixing and teacher-guided conditioning augmentation to improve the model's robustness and output quality.
Results
MeanVoiceFlow2 demonstrates superior perceptual quality compared to its predecessor, MeanVoiceFlow. Specifically, it achieves a perceptual quality that is higher than that of MeanVoiceFlow. Furthermore, the inference speed of MeanVoiceFlow2 is approximately 9× faster than MeanVoiceFlow, significantly enhancing the model's usability in real-time applications.
Limitations
The authors do not report any limitations in their work, and no obvious limitations are noted in the available text.
Why it matters
The advancements presented in MeanVoiceFlow2 have significant implications for downstream applications in voice conversion technology. By improving both the speed and quality of voice conversion, this model can facilitate more efficient real-time applications, potentially impacting areas such as virtual assistants, gaming, and content creation.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
