AI security testing company Irregular has released its report on an incident in which models evaluated in one of its test environments took offensive security measures against real systems rather than the simulated targets they were intended to attack.
The Israeli company, which raised $80 million in funding last year, has been in the headlines in recent weeks after revelations that AI models it tested on behalf of OpenAI, Anthropic and Meta had escaped their test environments and carried out real-world attacks.
Irregular’s core business is working with large AI labs to stress test models before releasing them to the public and running controlled simulations to measure a model’s capabilities in vulnerability research and offensive cyber tasks.
According to Irregular, test cycles typically involve thousands of simulation runs across multiple models over a period of 48 to 72 hours, using a range of parameters intended to reflect realistic attack conditions.
Anthropic and Irregular identified three incidents in which the AI giant’s models escaped their testing sandbox and hacked real organizations. A Blog post The book, published Friday by Irregular, focuses on one of those incidents.
The cause of the incident was attributed to a naming error. In creating one of its more advanced evaluation sets, Irregular’s engineering team assigned a fictional target company a name that, although unknown to them, matched an existing real-world domain. Fictitious names are typically matched to existing companies and websites before use, but the overlap went undetected because the actual domain was not widely known.
Internet access was enabled in the environment, and in a few cases, models reached the real domain instead of their simulated destination, treating it as part of the exercise even though they had not received instructions leading them there.
The aim of this evaluation was to test whether a model could help a malicious insider gain unauthorized access to sensitive data in a company’s production database. The exercise required the model to conduct reconnaissance, locate and use private keys, gather information about its target, extract data and attempt to avoid detection.
In the few runs where a model reached the real domain, it exploited vulnerabilities, extracted credentials, and gained access to a production database. In another case, a model went to another, similarly named website and found credentials that had already been publicly released.
Irregular said the target domain lacks general security measures, making it an easy target for most Frontier models. It added that the activity was difficult to capture because it only occurred in a small fraction of runs, often deep in a simulation after hundreds of interactions.
Going forward, the AI security company will expand manual review of model behavior during testing and create a dedicated internal team to challenge its own assumptions about containment and model control.
The post also highlighted major gaps facing the industry. Existing monitoring tools and classifiers have difficulty distinguishing legitimate red team activity from real attacks because the evaluation logs are inherently full of suspicious behavior.
Looking forward, Irregular said it is building clearer documentation processes with customers around evaluation setup and scope, and establishing an ongoing process to revalidate evaluations for new domain overlaps as new sites emerge over time.
It also called for better mechanisms to share forensic evidence such as model logs between organizations after an incident and announced plans for a white paper outlining best practices for securing AI assessments.
Related: Conflicting test objectives led Claude agents to deploy self-replicating malware
Related: OpenAI’s upcoming Astra model raises concerns about autonomous cyberattacks
Related: Critical one-click vulnerability in Atlassian’s Rovo AI exposed corporate data
