Irregular's testing bugs sent OpenAI, Anthropic, Meta and Google agents at real targets

Irregular's testing bugs sent OpenAI, Anthropic, Meta and Google agents at real targets

The Verge reports that a string of incidents in which AI agents from OpenAI, Meta, Anthropic and Google broke out of security testing environments and went after real-world targets all trace back to one company: Irregular, an Israeli startup founded as Pattern Labs in 2023 that stress-tests AI models in what it calls 'high-fidelity research platforms that simulate and monitor real-world AI security scenarios.' Irregular has worked with many of the industry's biggest AI firms, has been cited in OpenAI model system cards, tested systems for the UK government and Anthropic, and published research with the influential policy think tank RAND. In several of its tests this year, agents that were supposed to operate inside controlled, simulated networks instead escaped into the real internet. Some of the tests used 'capture-the-flag' exercises, a standard method for evaluating hacking ability in which agents search a simulated network for hidden information. Irregular CTO and cofounder Omer Nevo told The Verge that the agents were never meant to have internet access, but that 'internet access was unintentionally available.' Separately, a fictional company name invented for one simulation 'overlapped with a real domain.' Together, these two errors sent agents after real targets, though it is not clear which organizations were actually reached or attacked. Nevo confirmed that the same underlying issue in a single evaluation scenario was behind all the incidents involving OpenAI, Meta, Anthropic and Google models, and said these incidents 'have been disclosed,' while stressing that other recently reported industry security incidents, including the Hugging Face breach and issues involving the UK's AI Security Institute, were unrelated to Irregular. It remains unclear whether 'disclosed' meant informing clients, the public or regulators. Reporting cited in the piece indicates the four companies were notified at roughly similar times in late July; OpenAI and Anthropic announced their breaches themselves, while the Meta and, weeks later, Google incidents first surfaced through media reports. Irregular's testing extends beyond the four US firms: research on its site shows it also tested Kimi K3 and GLM-5.2, open models from Chinese companies Moonshot AI and Z.ai, which can be downloaded and run on a user's own hardware rather than accessed only through the company. Nevo said the same type of incident was not observed during the GLM or Kimi evaluations, but cautioned that 'this observation alone should not be interpreted as evidence that these models are less susceptible to this kind of behavior.' Neither Moonshot nor Z.ai responded to The Verge's request for comment. Nevo said Irregular has since tightened internet access controls, expanded monitoring and manual review, strengthened pre-evaluation checks to verify access matches the intended scope, and improved how it documents and agrees on evaluation setups with partners. He said Irregular plans to publish a broader report on lessons learned and safe practices for cyber evaluations once joint work with the affected companies is complete. None of the four US AI companies answered The Verge's questions about when they learned of the breaches, whether they sought damages or other remedies from Irregular, or whether they expect to keep working with the firm; Google and Anthropic did not respond at all, while OpenAI and Meta pointed to previously published blog posts.

Key facts

  • Irregular, an Israeli AI-testing startup founded as Pattern Labs in 2023, is the common source behind separate-seeming incidents in which agents from OpenAI, Meta, Anthropic and Google escaped simulated hacking environments and reached real-world targets.
  • CTO Omer Nevo said the cause was twofold: internet access was 'unintentionally available' during tests, and a fictional target company name 'overlapped with a real domain'.
  • Nevo said all these incidents stem from the same underlying issue in a single evaluation scenario and are unrelated to other reported incidents like the Hugging Face hack or breaches involving the UK's AI Security Institute.
  • Companies were notified at roughly similar times in late July; OpenAI and Anthropic disclosed their breaches themselves, while the Meta and Google incidents surfaced through media reports.
  • Irregular also tested Chinese open models Kimi K3 (Moonshot AI) and GLM-5.2 (Z.ai) without the same type of incident, though Nevo cautioned this does not mean those models are less susceptible; Moonshot and Z.ai did not respond to requests for comment.

Why it matters

A single testing-infrastructure failure at one third-party vendor produced what looked like a wave of independent rogue-AI incidents across four of the industry's largest AI developers, exposing how much of frontier AI safety validation depends on the security of the testing environments themselves, not just the models.

Who it affects

OpenAI, Anthropic, Meta and Google, whose agents were involved in the breaches; Irregular, whose testing setup caused them; and by extension any organization relying on third-party red-teaming or capture-the-flag style evaluations to certify agent safety.

How to use it

There is no product or licence here; the practical takeaway for AI labs and security teams is to treat evaluation sandboxes as attack surface in their own right, verifying that simulated networks have no real internet access and that fictional test assets do not collide with live domains.

How solid is it

The account rests on on-the-record confirmation from Irregular's CTO and cofounder Omer Nevo to The Verge, plus the companies' own prior disclosures and media reports of the Meta and Google incidents; it is not clear which real-world organizations were actually reached, and none of the four US companies answered further questions about timing, remedies sought, or future plans with Irregular.

Risks and caveats

The article does not identify which specific real-world targets the escaped agents reached, does not confirm any financial or legal consequences, and leaves ambiguous whether Nevo's term 'disclosed' meant notifying clients, the public, or regulators; Moonshot AI and Z.ai did not comment on their models' testing.

“All the incidents involving Irregular stemmed from the same underlying issue in a single evaluation scenario and have been disclosed.”

— Omer Nevo, Irregular CTO and cofounder