Notableefficiency inference

MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo

Published
Sep 30, 2026 — 16:31 UTC

Problem

The paper addresses a gap in the computational efficiency of content encoders in one-step voice conversion models. Existing models, such as MeanVoiceFlow, struggle with speed and quality, necessitating improvements for practical applications. This work is a preprint and has not undergone peer review.

Method

The authors introduce MeanVoiceFlow2, which employs a training method called conversion distillation. This method utilizes MeanVoiceFlow alongside real data reconstruction to enhance the model's performance. Additionally, the training incorporates techniques such as Diffusion-GAN training, which involves sample mixing and teacher-guided conditioning augmentation to improve the model's robustness and output quality.

Results

MeanVoiceFlow2 demonstrates superior perceptual quality compared to its predecessor, MeanVoiceFlow. Specifically, it achieves a perceptual quality that is higher than that of MeanVoiceFlow. Furthermore, the inference speed of MeanVoiceFlow2 is approximately 9× faster than MeanVoiceFlow, significantly enhancing the model's usability in real-time applications.

Limitations

The authors do not report any limitations in their work, and no obvious limitations are noted in the available text.

Why it matters

The advancements presented in MeanVoiceFlow2 have significant implications for downstream applications in voice conversion technology. By improving both the speed and quality of voice conversion, this model can facilitate more efficient real-time applications, potentially impacting areas such as virtual assistants, gaming, and content creation.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI