Armadin's AI swarm claims record live cyberattack test

AI agents are becoming as good at attacking organizations as at running them, The Register reports, and that is forcing security leaders to start hacking themselves before someone else does it for them. Matt Hartman, the former acting head of cyber at the US Cybersecurity and Infrastructure Security Agency (CISA) and now chief strategy officer of Merlin Group, warned that agentic AI creates a whole new class of non-human identities that are hard to manage and can slip past static security policies. He said organizations need to treat every agent as a privileged identity, since agents are moving from simply generating content to taking real actions with access to sensitive systems and data. Hartman also pointed to a surge in AI-amplified social engineering, including highly personalized phishing and convincing impersonation that make traditional trust signals unreliable, and said defenders need to double down on phishing-resistant authentication, behavioral signals and zero-trust principles.
Rob Joyce, the former NSA cyber boss, made the same point at a talk at the RSA Conference: "You are going to be red-teamed whether you pay for it or not. The only difference is, you know who gets the results delivered to them." Hartman agreed, saying a market is emerging for continuous, AI-native automated red-teaming and penetration testing, which he called a category every organization, including federal agencies, needs in the near term just to keep pace with attackers who can now find and exploit vulnerabilities in seconds rather than days.
The clearest example is Armadin, a new startup founded by Kevin Mandia, the founder and former CEO of Mandiant, which launched in March with $190 million in seed and Series A funding. Armadin builds autonomous attacker swarms, thousands of AI agents that run around the clock inside a client's infrastructure to simulate real-world attackers. Ahead of Black Hat this month, Armadin and agentic security operations provider Tenex.ai said they carried out what they called the largest controlled live AI cyberattack on record against an unnamed "leading" global institution. Over three days, Armadin's swarm generated 17 million offensive actions, discovered 38 validated attack paths and produced 238 security findings, while Tenex.ai's platform triaged 100 percent of 101,169 alerts and reconstructed the entire attack across 231 billion raw events. The two companies estimate the same exercise would have taken a five-person human analyst team about 2,400 hours, or four months, to complete.
Armadin co-founder and chief offensive security officer Evan Pena, who previously led Mandiant's 210-person global red team, said one lesson from OpenAI's models autonomously attacking Hugging Face is that organizations need to run safe offensive AI tests against their own systems, since OpenAI's rogue models in that case deliberately had no guardrails. Pena said AI agents give defenders three things human-only teams never had: unlimited time, since agents do not sleep or take holidays; pre-trained and post-trained expertise in code review, application security and network exploitation; and full coverage, since a team that could once check 1,000 to 2,000 of an organization's 10,000 external systems in a given period can now check all 10,000 in hours. He said Armadin's agents have broken into every single customer's environment and have found more than 50 zero-days that allow remote code execution on live systems, which he called high-impact rather than cosmetic bugs.
Jay Bavisi, founder and group president of EC-Council, told The Register that even the best organizations only pen-test once a year for compliance, or quarterly at best, because human-led pen-tests take about three months. That leaves a speed problem, since attackers test continuously; a scope problem, since human testers rarely check an entire organization; and a sophistication problem, since results vary between individual testers. In June, EC-Council began offering pen-testing professionals a sponsored attempt at its CPENT AI exam: the council donates $1,000 in training and certification credits to nonprofit partners for every participant who passes, and $250 for every completed training program regardless of outcome, up to a $1 million cap. Bavisi said the traditional once-a-year or once-a-quarter pen-testing model is going away in favor of automated testing, but argued the pen-tester role will not disappear. It will evolve toward understanding business impact, testing the robustness of AI systems themselves and defining guardrails for agentic behavior.
Key facts
- Armadin, founded by ex-Mandiant CEO Kevin Mandia, launched in March with $190 million in seed and Series A funding to build autonomous AI attacker swarms.
- Armadin and Tenex.ai say they ran the largest controlled live AI cyberattack on record: over three days, Armadin's swarm generated 17 million offensive actions, found 38 validated attack paths and produced 238 findings against an unnamed global institution.
- Tenex.ai's platform triaged 100 percent of 101,169 alerts and reconstructed the attack across 231 billion raw events; the companies say the exercise would have taken a five-person analyst team about 2,400 hours, or four months, by hand.
- Armadin co-founder Evan Pena says the firm's agents have broken into every customer's environment and found over 50 zero-days that allow remote code execution.
- EC-Council's Jay Bavisi says human-led pen-tests take about three months, forcing annual or quarterly cadence, and the group is running a sponsored CPENT AI exam program (up to $1,000 per pass, $250 per completed training, $1 million cap) since June.
Why it matters
AI agents cut both ways: attackers already use them to automate reconnaissance and exploit discovery at machine speed, and Rob Joyce's warning that any organization will be red-teamed whether it pays for it or not is the crux of the piece. Matt Hartman describes a burgeoning market for continuous, AI-native red-teaming and pen-testing, and Armadin's exercise is the largest public demonstration to date of what that looks like: a live attack simulation that ran in three days versus the roughly 2,400 hours, or four months, a five-person human team would have needed.
Who it affects
Enterprises and federal agencies running agentic AI and managing the resulting non-human identities; security and pen-testing teams whose work is being automated and reshaped; vendors like Armadin, Tenex.ai and Merlin Group's portfolio companies building the tooling; and the unnamed global institution that was the actual target of Armadin's live-attack exercise.
How to use it
Hartman's concrete guidance: treat every AI agent as a privileged identity, invest in phishing-resistant authentication, behavioral signals and zero-trust controls. On the offense side, Bavisi points to EC-Council's CPENT AI certification, with a sponsored exam attempt and training-credit donations to nonprofits running since June, as a way for pen-testers to reskill toward AI-native testing rather than being displaced by it.
How solid is it
The exercise's scale is reported directly by Armadin and Tenex.ai about their own engagement, and The Register notes Evan Pena's framing is self-serving since it promotes Armadin's own business. The target institution is not named, so the numbers cannot be independently checked against a public record, though named, credentialed sources at CISA, the NSA, Mandiant's former leadership and EC-Council are quoted directly and consistently on the broader trend.
Risks and caveats
The 190 million dollar figure is funding, not a valuation, and Armadin's own claim to the "largest" live AI cyberattack on record is unverified against any independent benchmark. The source does not say which year Armadin launched in March, does not name the targeted institution, does not date Rob Joyce's RSAC talk beyond "during a talk," and does not say whether or how the more than 50 zero-days Armadin found were disclosed or patched.
“You are going to be red-teamed whether you pay for it or not. The only difference is, you know who gets the results delivered to them.”
— Rob Joyce, former NSA cyber boss