DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
Zhengming Yu, Junkun Yuan, Haotian Yang, Gordon Guocheng Qian, Yizhi Wang, Angtian Wang, Yiding Yang, Bo Liu, Xin Li, Wenping Wang, Chongyang Ma
- Published
- Oct 1, 2026 — 17:59 UTC
Problem
The paper addresses the challenge of efficiently training a student model for visual generation tasks without relying on auxiliary score fitting. This is particularly relevant in the context of generative models, where traditional methods often require complex training regimes that can be computationally expensive and time-consuming. The authors propose a new approach, Distribution Matching as Adversarial Distillation (DMAD), to streamline this process. Notably, this work is presented as a preprint and has not undergone peer review.
Method
The core technical contribution of this paper is the DMAD algorithm, which employs a unique architecture featuring two discriminator heads on a shared backbone. The loss function utilized is based on linear losses applied to the discriminator logits, facilitating effective training of the student model. The training data comprises three datasets: ImageNet-64x64, COCO-10K, and MiniMax-H3-33B. An additional mechanism, termed gap-based reweighting, is introduced to adapt teacher supervision across varying noise levels, enhancing the robustness of the training process. However, the authors do not specify the training compute resources utilized in their experiments.
Results
The results demonstrate the effectiveness of the DMAD approach across several benchmarks:
- On ImageNet-64x64, the model achieves a Fréchet Inception Distance (FID) of 1.04 for one-step generation.
- For COCO-10K, the FID is reported as 14.47 for four-step SDXL generation.
- The VBench total score is 85.15, evaluated against the four-step Wan2.1-T2V-14B model.
- Human preference rates indicate a preference of 79.1% for DMAD over DMD2 and 84.6% over rCM, both evaluated on the four-step student model using MiniMax-H3-33B.
The available text does not report quantitative results for baseline comparisons.
Limitations
The authors do not report any limitations in their work. However, the absence of specified training compute resources may hinder reproducibility and scalability assessments. Additionally, the lack of baseline comparisons for some results could limit the contextual understanding of the performance improvements.
Why it matters
The implications of this work are significant for downstream applications in visual generation, particularly in scenarios where computational efficiency is paramount. By eliminating the need for auxiliary score fitting, DMAD could facilitate faster training and deployment of generative models, making them more accessible for real-time applications. This approach may also inspire further research into adversarial distillation techniques and their potential to enhance model performance in various domains.
By Turing Wire Research Desk · Oct 1, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
