Majorsafety alignmentOpenAI

OpenAI Halts Model Reasoning Theft; Azure Attacks Persisted

Published
Oct 1, 2026 — 12:05 UTC

OpenAI Halts Model Reasoning Theft; Azure Attacks Persisted

On July 1, 2026, a campaign to steal reasoning from AI models began, ultimately resulting in over 16,000 attack requests from more than 4,000 users between July 24-25. OpenAI responded by shutting down over 15,000 accounts linked to these attacks on July 28. Despite these efforts, Joachim Schaeffer, a researcher involved in the attacks, noted that the same models exhibited varying levels of protection depending on the platform, stating, "Same models, but different protections depending on which platform serves them."

OpenAI's GPT-6 Astra and Opus 4.8 were among the models targeted, while Anthropic's Sonnet 5 also faced similar extraction attempts. Following a series of incidents, OpenAI successfully blocked the attacks on its and Anthropic's APIs by September 13, 2026. On September 27, OpenAI implemented additional safeguards for the Azure endpoint, leading to a cessation of successful extractions from Anthropic models on Azure by September 28.

Despite OpenAI's measures, Schaeffer claimed, "A cheap model can unlock an expensive model's hidden thoughts," indicating that vulnerabilities remain. OpenAI acknowledged that the issue extends beyond its models, emphasizing the need for comprehensive patches that address all types of attacks across various cloud platforms. This follows previous coverage of AI model security challenges, underscoring the ongoing risks in cloud-based AI deployments.

Summarised from The Decoder's original report by the Turing Wire Newsdesk. Read the original for the full story.

Source: The Decoder