HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment
Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran
- Published
- Sep 29, 2026 — 16:56 UTC
Problem
Local language models often lag behind frontier models in capability, particularly when handling complex queries. This limitation can undermine the privacy and cost benefits associated with local deployment. The paper addresses this gap by proposing a method that enhances the performance of local models during inference, particularly in challenging scenarios. Notably, this work is presented as a preprint and has not undergone peer review.
Method
HARISSA fine-tunes a language model to improve its inference capabilities. The core algorithm leverages hidden states to predict the correctness of generated answers, allowing for a more nuanced decision-making process. The decision mechanism cascades through various answering methods based on both prefill and answer states, optimizing the response generation process. Specific details regarding the loss function, data used for training, and training compute resources are not disclosed in the paper.
Results
HARISSA achieves an accuracy that is within one point of the chain-of-thought method while demonstrating a latency that is 2.7 times lower. In comparative evaluations, HARISSA outperforms FrugalGPT and Self-REF cascades in terms of accuracy at equivalent latency levels. Additionally, it exhibits a lower deferral rate, producing fewer incorrect answers than standard confidence signals across five out of six task and setting pairs. The available text does not report quantitative results beyond these comparisons.
Limitations
The authors do not report any limitations in their work. However, the lack of specified details regarding the loss function, data, and training compute may hinder reproducibility and further evaluation of the method's robustness.
Why it matters
The implications of HARISSA are significant for the deployment of local language models, particularly in applications where privacy and efficiency are paramount. By improving the accuracy and reducing latency, HARISSA enables more reliable and faster responses in real-time applications, potentially expanding the use cases for local models in sensitive environments.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
