Notableefficiency inferenceNVIDIA

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica

Published
Sep 29, 2026 — 17:59 UTC

Problem

The paper addresses a significant gap in the efficient processing of long-context sequences in large language models (LLMs), specifically focusing on the inference bottleneck caused by state updates. This issue is particularly relevant in scenarios where maintaining high performance while managing extensive context is critical. The authors propose a solution to this problem through a novel quantization technique, which is essential for improving computational efficiency without sacrificing accuracy. Notably, this work is presented as a preprint and has not undergone peer review.

Method

The proposed method, named LeapQuant, employs an innovative approach to quantization, specifically targeting the recurrent states of the model. Key components of the method include:

  • 8-bit Recurrent-State Quantization: This technique reduces the precision of the recurrent states to 8 bits, significantly lowering memory and computational requirements.
  • Per-window Quantization: The quantization process occurs at the end of a token window, allowing for efficient state updates while processing sequences.
  • Compensator Tokens: High-precision tokens are retained for the largest outliers, ensuring that critical information is preserved despite the quantization.
  • Error Mitigation: The authors implement a smoothing technique for residuals prior to quantization, which helps to minimize the impact of quantization errors.
  • Training-Free Method: LeapQuant does not require additional training, making it a practical solution for existing models.

Results

The results demonstrate significant improvements in computational efficiency:

  • Kernel Speedup: LeapQuant achieves a speedup of 2.05 to 3.70 times compared to the FP32 baseline.
  • End-to-End Inference Speedup: The method provides a 1.47 times speedup in end-to-end inference relative to the FP32 baseline.
  • Accuracy: The accuracy of LeapQuant remains comparable to the FP32 baseline, indicating that the quantization does not adversely affect model performance.

Limitations

The authors acknowledge potential limitations, including:

  • Model Quality Degradation: There is a risk of degradation in model quality due to rounding errors associated with the quantization process. The handling of outliers is not explicitly addressed, which could lead to further inaccuracies.

Why it matters

The implications of LeapQuant are significant for downstream applications that require efficient long-context processing in LLMs. By providing a method that balances computational efficiency with accuracy, this work opens avenues for deploying LLMs in resource-constrained environments. Furthermore, the training-free aspect of LeapQuant allows for easy integration into existing models, potentially accelerating the adoption of efficient inference techniques in practical applications.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI