LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica
- Published
- Sep 29, 2026 — 17:59 UTC
Problem
The paper addresses a significant gap in the efficient processing of long-context sequences in large language models (LLMs), specifically focusing on the inference bottleneck caused by state updates. This issue is particularly relevant in scenarios where maintaining high performance while managing extensive context is critical. The authors propose a solution to this problem through a novel quantization technique, which is essential for improving computational efficiency without sacrificing accuracy. Notably, this work is presented as a preprint and has not undergone peer review.
Method
The proposed method, named LeapQuant, employs an innovative approach to quantization, specifically targeting the recurrent states of the model. Key components of the method include:
- 8-bit Recurrent-State Quantization: This technique reduces the precision of the recurrent states to 8 bits, significantly lowering memory and computational requirements.
- Per-window Quantization: The quantization process occurs at the end of a token window, allowing for efficient state updates while processing sequences.
- Compensator Tokens: High-precision tokens are retained for the largest outliers, ensuring that critical information is preserved despite the quantization.
- Error Mitigation: The authors implement a smoothing technique for residuals prior to quantization, which helps to minimize the impact of quantization errors.
- Training-Free Method: LeapQuant does not require additional training, making it a practical solution for existing models.
Results
The results demonstrate significant improvements in computational efficiency:
- Kernel Speedup: LeapQuant achieves a speedup of 2.05 to 3.70 times compared to the FP32 baseline.
- End-to-End Inference Speedup: The method provides a 1.47 times speedup in end-to-end inference relative to the FP32 baseline.
- Accuracy: The accuracy of LeapQuant remains comparable to the FP32 baseline, indicating that the quantization does not adversely affect model performance.
Limitations
The authors acknowledge potential limitations, including:
- Model Quality Degradation: There is a risk of degradation in model quality due to rounding errors associated with the quantization process. The handling of outliers is not explicitly addressed, which could lead to further inaccuracies.
Why it matters
The implications of LeapQuant are significant for downstream applications that require efficient long-context processing in LLMs. By providing a method that balances computational efficiency with accuracy, this work opens avenues for deploying LLMs in resource-constrained environments. Furthermore, the training-free aspect of LeapQuant allows for easy integration into existing models, potentially accelerating the adoption of efficient inference techniques in practical applications.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
