Notableefficiency inference

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Sahil Mahendrakar

Published
Oct 5, 2026 — 17:56 UTC

Problem

This work addresses the challenge of distilling a large text-to-speech (TTS) model, Kokoro-82M, into a significantly smaller model, Paradee, with only 8 million parameters. The motivation stems from the need for efficient TTS systems that maintain performance while reducing computational requirements. The paper is a preprint and has not undergone peer review.

Method

  • Model Architecture: Paradee retains the architectural framework of Kokoro-82M but implements narrower layers to reduce parameter count to 8.07 million.
  • Training Method: The distillation process involves training two halves of the model separately against a frozen teacher model (Kokoro-82M), allowing for effective knowledge transfer.
  • Data: The training corpus is synthesized from the teacher model, ensuring that key features such as durations, pitch, energy, and phoneme characteristics are preserved.
  • Training Compute: Paradee requires 15 times less computational resources compared to the original Kokoro-82M model, enhancing its accessibility for deployment.
  • Loss Function: The training employs spectral losses followed by adversarial training to improve the quality of the generated speech.
  • Quantization: Model weights are quantized to int8, optimizing memory usage and inference speed.
  • Alignment Learning: The model does not require alignment learning, simplifying the training process.
  • Joint Training: No joint training is necessary, further streamlining the distillation approach.
  • Deployment: Paradee is designed to run efficiently on a single laptop, making it suitable for edge applications.

Results

  • Model Size: Paradee has a compact size of 8.5 MB, facilitating easier deployment.
  • Inference Speed: The model achieves an inference speed that is 25 times faster than real-time on a single CPU thread, indicating high efficiency.
  • UTMOS Score: Paradee scores 4.41 on the UTMOS benchmark, compared to the teacher model's score of 4.52, demonstrating competitive performance despite the reduction in size.

Limitations

The authors note that the initial version of Paradee exhibited a slight buzzing sound in the audio output, particularly in the phase of voiced speech within the frequency range of 2 to 8 kHz. This artifact may affect the perceived quality of the synthesized speech. Additionally, the paper does not discuss potential limitations related to the generalization of the model across diverse datasets or languages.

Why it matters

The development of Paradee signifies a substantial advancement in the field of TTS by demonstrating that effective distillation can yield a lightweight model without significant loss in performance. This has implications for deploying TTS systems in resource-constrained environments, such as mobile devices or embedded systems, and opens avenues for further research into efficient model compression techniques in deep learning.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI