MCP trust gaps let one compromised AI agent pass attacks to others

AI agents are now used in millions of organizations, and Ars Technica reports that this is opening new ways for attackers to make them act maliciously, for example by pulling out database contents and sensitive business and personal information.
In the past five months, Google and four other organizations, which the article says have little in common beyond their use of AI agents, have acknowledged vulnerabilities of one kind. An attacker exploits one agent inside a targeted network and uses it to spread harmful instructions to other internal agents. The technique is a special form of prompt injection. It targets not the LLM but a particular agent, such as one built for translation or data analysis. Guardrails inside such agents, if they exist at all, are often lax and pass the instructions along to other agents down the chain. Because the downstream agent explicitly trusts the first one, it follows the directions.
Independent researcher Syed Anas Mohiuddin tested agents from organizations including Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government. His proof-of-concept attacks exploit trust gaps in MCP, short for Model Context Protocol. The standard is one way AI apps and agents communicate with each other inside an internal network.
The article gives two reasons the attacks work. Many special-purpose agents lack the guardrails that might normally limit the worst effects of a prompt injection. And MCP servers store credentials for each agent, while agents are built to trust every other internal agent, so an exploit that the LLM would have rejected succeeds. In many cases, well-crafted prompts aimed at the right agent lead to server-side request forgery, a vulnerability that makes a web server send unauthorized network requests.
The piece's own subheading calls the problem unexpected and hard to mitigate.
Key facts
- Google and four other organizations have acknowledged, in the past five months, vulnerabilities where one compromised internal AI agent passes malicious instructions to other internal agents.
- Independent researcher Syed Anas Mohiuddin built proof-of-concept attacks that exploit trust gaps in the Model Context Protocol (MCP) and tested agents from Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government.
- The technique is a form of prompt injection aimed at a specific agent, such as one for translation or data analysis, rather than at the LLM.
- MCP servers store credentials for each agent and agents are built to trust every other internal agent, so an exploit the LLM would have rejected succeeds.
- In many cases, well-crafted prompts aimed at the right agent lead to server-side request forgery.
Why it matters
Most prompt-injection defences focus on the LLM. This attack goes around it: it targets an individual agent, and the next agent in the chain accepts the instructions because it is built to trust its peers. The article says that is why an exploit the LLM would have rejected can still succeed. With AI agents adopted in millions of organizations, per the article, a weakness in how agents trust each other is a structural issue rather than one product's bug.
Who it affects
Organizations that run several AI agents inside one internal network and connect them over MCP. Mohiuddin tested agents from Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government. Separately, Google and four other organizations have acknowledged vulnerabilities of this type in the past five months. The article does not say which of the listed organizations are the ones that acknowledged them.
How to use it
This is a research finding, not a tool, and the available text gives no recommended mitigations. The mechanics it describes point to two places to look in any multi-agent setup: whether special-purpose agents have guardrails at all, and how much one agent trusts instructions arriving from another over MCP.
How solid is it
The reporting comes from Ars Technica and rests on proof-of-concept attacks by one independent researcher, Syed Anas Mohiuddin. Google and four other organizations have acknowledged vulnerabilities of this kind. The available text names no CVE identifiers, severity scores or patch details, and quotes no one from the researcher or the organizations. The original article is only partly available, so details such as the specific agents tested are not covered here.
Risks and caveats
No evidence of exploitation in the wild is stated; the attacks described are proof-of-concept. The article says server-side request forgery follows in many cases, not all. Because the weakness lies in the trust between agents and the credentials MCP servers hold for each of them, the article presents it as hard to mitigate, and no fix is described.
“Because the latter agent explicitly trusts the first one, it follows the directions.”
— Ars Technica