OpenAI reasoning traces: Company disrupted an extraction campaign
OpenAI said it disrupted coordinated activity aimed at obtaining protected internal reasoning traces from its models. The company denied any compromise of databases or access to users’ stored conversations.

On September 30, OpenAI said it had identified and disrupted a coordinated campaign aimed at extracting protected internal reasoning traces from models, also referred to as OpenAI reasoning traces. According to the company, this did not involve breaking encryption, compromising a database, or directly accessing users’ stored conversations.
Reasoning traces are internal intermediate steps that a model uses when solving tasks. OpenAI protects them in part because obtaining them could make it easier to imitate a model’s capabilities or bypass security measures without comparable development and security costs.
OpenAI reasoning traces and the campaign timeline
According to OpenAI, the activity began on July 1. On July 24 and 25, the company recorded approximately 16,000 attempted requests from more than 4,000 accounts. By July 28, it had disrupted the related cluster of more than 15,000 accounts.
The published figures describe extraction attempts, not necessarily the successful acquisition of protected traces. OpenAI did not disclose how much data the attackers may have obtained, or whether it could have been used to train another model.
The company said it deployed account blocks or restrictions against the activity and strengthened controls during registration and at the infrastructure level. It also closed one path that enabled replay of encrypted reasoning and expanded detection of potential leakage. According to the company, the investigation and further protective measures are ongoing.
Researchers have already described a similar class of attacks
In August, independent researchers described a class of attacks based on reusing encrypted reasoning blocks between sessions or models. After responsibly reporting the vulnerabilities, they presented demonstrations involving systems from OpenAI, Anthropic, and Google.
This research does not independently confirm the specific campaign announced by OpenAI. It does, however, confirm that reusable encrypted blocks have been technically investigated and that model providers must address the protection of internal outputs beyond the usual security of user accounts.
Attribution has not yet been independently confirmed
OpenAI attributed the core of the activity to individuals linked to Moonshot AI. However, this is OpenAI’s own assessment; the available materials do not independently confirm the attribution. Moonshot AI’s response and further technical details about the specific attempts are not included in the published information.
As developments continue, it will be important to see whether independent confirmation of the attribution or a statement from Moonshot AI emerges. Attention will also focus on whether other AI model operators report similar attempts, and whether OpenAI publishes more information about the effectiveness of the measures introduced and any potentially successful extractions.
Sources
- OpenAI – Confirms the campaign announcement, timeline, scope of the attempts, description of mitigations, and OpenAI’s attribution assessment.
- arXiv – Stealing Reasoning Traces from Proprietary LLM APIs – Documents a class of attacks on reusable encrypted reasoning blocks described by researchers and proposed mitigations; it does not independently confirm the specific campaign.
Verified and updated: 10/01/2026 06:26



