Notableefficiency inference

MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

Jiarui Chen, Zeqiang Lai, Jiangshan Wang, Ziheng Ouyang, Ye Huang, Xiangyu Yue, Cewu Lu, Chunchao Guo

Published
Oct 5, 2026 — 17:51 UTC

Problem

Existing sparse attention methods in diffusion transformers suffer from a degradation in generation quality and fidelity when high sparsity levels are employed. This paper addresses this gap by proposing a new approach that maintains performance while improving computational efficiency. The work is presented as a preprint and has not yet undergone peer review.

Method

The authors introduce the Meta-Cached Sparse Attention (MC-Sparse) framework, which innovatively selects individual key-value (KV) tokens and organizes similar queries into tile-aligned groups. The method utilizes metadata caching, where query groups and KV indices are selected based on exact attention probabilities and the residuals between dense and sparse attention outputs. This approach allows for efficient execution on GPUs, enabling the reuse of metadata across subsequent denoising steps, thus optimizing the attention mechanism in diffusion models.

Results

The proposed MC-Sparse framework demonstrates significant improvements in denoising speed compared to traditional dense attention methods. Specifically, it achieves a speedup of 1.80x on the Minimax-H3-Base benchmark and a 2.32x speedup on 3D asset generation tasks, showcasing its effectiveness in enhancing computational efficiency without compromising output quality.

Limitations

The authors do not report any limitations in their work, and no obvious limitations are identified in the available text.

Why it matters

The implications of this research are substantial for downstream applications in generative modeling, particularly in scenarios where computational resources are constrained. By closing the dense-sparse attention gap, MC-Sparse paves the way for more efficient training and inference in diffusion transformers, potentially leading to broader adoption and improved performance in various AI tasks.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI