KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
- Published
- Oct 1, 2026 — 17:59 UTC
Problem
This work addresses a gap in the evaluation of large language models (LLMs) regarding their ability to generate executable commands for cybersecurity tools. The authors highlight the lack of fine-grained benchmarks that can assess the performance of LLMs in this specific domain, particularly in generating commands that are not only syntactically correct but also executable in a real-world context. The paper is a preprint and has not yet undergone peer review.
Method
The authors present KaliBench, a benchmark consisting of 8,504 query-command pairs across 1,642 tools, categorized into 23 capability dimensions and 5 security phases. The construction of this benchmark employs a manuscript-grounded pipeline that includes deterministic canonicalization and alias-aware evaluation to ensure the accuracy and relevance of the command generation.
A multi-stage verification pipeline is introduced, which combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement to verify the generated commands. This approach allows for runtime-free verifiable rewards to be utilized during the training process, enhancing the model's ability to generate accurate commands without the need for real-time execution during training.
Results
The benchmark reveals that the exact-command accuracy achieved is 42% in an unrestricted setting. Furthermore, the authors report that employing supervised fine-tuning and reinforcement learning with verifiable rewards significantly improves the performance of an 8B parameter model, bringing it to a level comparable to that of a 685B mixture of experts (MoE) model.
Limitations
The authors note that no open-weight model currently exceeds the 42% exact-command accuracy in the unrestricted setting. This limitation suggests that while the benchmark provides valuable insights, there is still significant room for improvement in the command generation capabilities of existing models.
Why it matters
KaliBench serves as a critical tool for advancing research in the intersection of natural language processing and cybersecurity. By providing a structured framework for evaluating LLMs in generating executable commands, it lays the groundwork for future developments in automated cybersecurity tools. The implications of this work extend to enhancing the reliability and effectiveness of LLMs in practical cybersecurity applications, potentially leading to more robust automated systems for threat detection and response.
By Turing Wire Research Desk · Oct 1, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
