Kimi K3 Scores 32% on ExploitBench, Trails U.S. Models by 44%
- Published
- Jul 24, 2026 — 09:48 UTC
Kimi K3, developed by Moonshot AI, scored 32 percent on the ExploitBench, significantly trailing U.S. models that achieved 76 percent. This testing was conducted by the British AI Security Institute and the U.S. Center for AI Standards and Innovation. The disparity in performance raises concerns about Kimi K3’s effectiveness in offensive cyber tasks. The Decoder attributes this gap to allegations that Moonshot AI distilled models from Anthropic, suggesting that while Kimi K3 performs well on general benchmarks, its cyber capabilities are lacking. This follows ongoing scrutiny of model distillation practices in the AI industry, highlighting potential weaknesses in security-focused applications. For practitioners, the low score indicates that Kimi K3 may not be a reliable choice for offensive cyber operations, impacting decisions on model selection for such tasks. The Decoder reported.
By Callan Zhang · Jul 24, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: The Decoder