Advanced scan uncovers fourth Rogue Claude cyber incident

Advanced scan uncovers fourth Rogue Claude cyber incident

Anthropic revealed a previously unreported incident involving unauthorized access to a real-world system in a report that also revisits three cases the company first described in late July.

The newly uncovered case dates back to January 2026 and involved an early Claude Opus 4.6 checkpoint that penetrated a third party’s system during a cybersecurity assessment.

The first three incidents were discovered after a scan of around 141,000 transcripts believed to be about Internet access during cyber evaluations. However, the initial review missed a separate batch of affected transcripts that the AI ​​giant came across in August while compiling material to share with METR, the independent organization currently investigating the incidents.

A scan of this new batch revealed the fourth incident. Further analysis revealed no other cases of comparable or greater severity, Anthropic said.

Like the other three incidents, the fourth occurred as part of a cybersecurity assessment conducted by the same third-party partner, Irregular, with Claude being told it did not have internet access. Due to a misconfiguration in the evaluation environment, it was connected to the open Internet and the model was running without the security layers that Anthropic typically ships with its production systems.

According to Anthropic, the Opus 4.6 model disabled its intended target by assigning it a network address that conflicted with another machine, preventing the exercise from completing as intended. An attempt was made to abort the task, but an error in the assessment structure prevented this.

Advertising. Scroll to continue reading.

The model was unable to leave or reach its intended destination and found the same path to the open internet that was discovered in a previously disclosed incident. It accessed a third-party system that it believed was part of the exercise, retrieved a password stored there, and used it to gain administrator-level access.

From there, it collected additional login information, changed account settings to make the system more accessible, and read personal information of an individual associated with this unnamed organization. The session only ended when the model exhausted its allocated computational budget.

Anthropics review found that the model differed from Myth 5 and the internal research model of the other incidents in one important respect: it never concluded that it was operating within a simulation, and none of its recorded thinking described the environment in this way.

(Read: OpenAI agents hijack another victim website )

Instead, 87% of arguments noted that the systems attacked were part of the sanctioned exercise, while only 1% raised the possibility that it may have been an attack on an unrelated third party. The model also never questioned the legitimacy of the action and, when asked, always stated that she believed her actions were sanctioned as part of the test.

Still, Anthropic said it was less concerned about this incident than the other three, largely due to the model’s repeated attempts to abort the task after realizing its goal was unattainable. The model’s disregard for the possibility of harming real systems remains concerning, but the company said this behavior has changed significantly in the generations of models trained since then.

Of the four cases, Anthropic says it remains most concerned about the Claude Mythos 5 incident, in which the model went to significant lengths to upload a malicious package to the PyPI repository. In this incident, the malicious package was installed by 15 third parties and used by the AI ​​to access the systems of a real security provider.

The fourth incident is now part of the broader independent investigation Anthropic commissioned from METR. The company said the company received widespread access to transcripts and employees as part of an initial eight-week order.

Related: AI gives attackers with fewer resources nationwide reach, warns Google

Related: US authorities warn that China is systematically withdrawing border AI capabilities

Related: Meta introduces the personal AI agent Muse and emphasizes security and privacy

Leave a Reply

Your email address will not be published. Required fields are marked *