The adoption of AI agents in millions of organizations is creating new opportunities for attackers to exfiltrate database contents and sensitive business and personal information. In the past five months, Google and four other organizations have acknowledged vulnerabilities that exploit one agent inside a network to spread harmful instructions to other internal agents.
Unexpected and hard to mitigate
The technique is a form of prompt injection that targets a particular agent, such as one for translation or data analysis, rather than the underlying LLM. Guardrails inside such agents are often lax, and because subsequent agents in a chain explicitly trust the first one, they follow the malicious directions. Independent researcher Syed Anas Mohiuddin tested agents from organizations including Google, JP Morgan Chase, Rapid7, and the US and French governments, finding trust gaps in the Model Context Protocol (MCP).
MCP is a standard used for communication between AI apps and agents inside an internal network. Many special-purpose agents lack the guardrails that might normally mitigate prompt injection. Since MCP servers store credentials for each agent, an exploit that would have been rejected by the LLM can succeed. In many cases, well-crafted prompts targeting the right agent lead to server-side request forgery (SSRF), a vulnerability that causes a web server to make unauthorized network requests.
Protocol Pivoting
Mohiuddin calls this class of attack protocol pivoting. It occurs when an adversary gains initial access through one protocol, exploits trust assumptions, and escalates to capabilities accessible via a different protocol, such as Google’s Agent-to-Agent (A2A) protocol. However, Markus Vervier, a researcher at X41 D-Sec, suggests the term indirect prompt injection remains more accurate, noting that the use of different protocols is not strictly required for such attacks to work.
The vulnerability affecting Google carried a severity rating of 8 out of 10. It stemmed from an MCP toolbox for databases that initialized its HTTP client without a CheckRedirect policy and failed to validate target IP addresses. Google’s fix involved applying an allow-list of IP ranges and block lists. In contrast, CVE-2026-97228, a vulnerability found in Rapid7’s network, carried a severity rating of 2.7 and was fixed last month.
Abandoning Zero Trust
The fact that this technique worked across five diverse organizations suggests that in the rush to build agentic architectures, many have abandoned the core security principle of zero trust. Under this model, nodes should require authorization before conducting sensitive transactions with others. "AI agents give attackers a fresh set of connections to walk across," Douglas McKee, director of vulnerability intelligence at Rapid7, told Ars. "Each protocol was built assuming it lived on its own, so each one checks its own front door while nobody watches the hallway in between."