STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei
- Published
- Sep 29, 2026 — 17:59 UTC
Problem
The paper addresses a gap in the quantization of recurrent states, specifically focusing on maintaining accuracy while reducing memory usage. The authors highlight the challenge of quantizing recurrent states without degrading performance, which is particularly relevant in the context of deploying large language models. This work is presented as a preprint and has not yet undergone peer review.
Method
The proposed framework, STEPQuant, employs a spatial-temporal post-training quantization approach tailored for Delta-rule recurrent states. Key components of the method include:
- Precision Allocation: The quantization process allocates precision based on the magnitude of errors and the memory lifetime of the states.
- Key-row and Value-column Fitting: The framework jointly fits the key-row and value-column based on the distributions of the states and the impact of key-rows on output errors.
- Models Used: The method is evaluated on models such as Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct.
- Budget: The quantization is designed to operate under a nominal 6-bit budget while matching the accuracy of full precision (FP32) states.
- Compression and Memory Reduction: STEPQuant achieves over 5x compression of recurrent states and reduces total serving memory by up to 68.7%.
- Implementation: The framework is integrated into SGLang, utilizing optimized GPU kernels for efficient execution.
Results
The results indicate that STEPQuant successfully matches the accuracy of FP32 states while operating under a nominal 6-bit budget. In comparison to a uniform INT8 configuration at 4 bits, STEPQuant demonstrates superior performance in terms of accuracy retention. Additionally, the method achieves a significant reduction in total serving memory, with a reported decrease of up to 68.7%. However, no specific baseline for memory reduction is provided in the text.
Limitations
The authors do not report any limitations in the study. However, the absence of a comparative baseline for memory reduction could be seen as a potential oversight, as it limits the contextual understanding of the memory savings achieved.
Why it matters
The implications of this work are significant for the deployment of large-scale recurrent models in resource-constrained environments. By enabling effective quantization without sacrificing accuracy, STEPQuant paves the way for more efficient model serving and potentially broader accessibility of advanced AI systems. This approach could influence future research in model compression and optimization, particularly in the context of real-time applications where memory and computational efficiency are critical.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
