Researchers have shown that the Model Context Protocol (MCP), the one‑way channel many organizations use for internal AI agent communication, can be hijacked to forward malicious prompts from one agent to another. The finding matters because any environment that strings together translation, analysis, or other specialized agents now faces a vector that bypasses traditional LLM‑level defenses.
What changed in agent‑to‑agent communication
Over the past five months, Google and four additional organizations disclosed vulnerabilities where a compromised agent injected a harmful instruction that was automatically relayed to downstream agents via MCP. The attack does not target the language model itself; instead it exploits the trust relationship embedded in the protocol. Guardrails inside the first agent, when present, were insufficient to stop the instruction from propagating.
Model Context Protocol trust gaps
Independent researcher Syed Anas Mohiuddin demonstrated proof‑of‑concept attacks against agents deployed by Google, JP Morgan Chase, Weviate, Rapid7, the French interministerial digital directorate, and the US federal government. The common factor was the use of MCP to pass context between agents. Because MCP is defined as a one‑way flow, the receiving agent assumes the incoming prompt is authorized, creating a trust gap that can be weaponized.
Architectural and operational implications
- Design assumptions: Systems that chain agents together should no longer assume that downstream agents are insulated from upstream prompt manipulation.
- Guardrail placement: Relying solely on LLM‑level filters is insufficient; each agent needs its own validation of incoming instructions before forwarding.
- Observability: Logging and tracing of MCP payloads become critical for detecting unexpected instruction chains.
- Deployment scope: Environments that expose MCP across multiple services or tenants must treat the protocol as a potential attack surface.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Teams should audit any use of MCP, verify that each agent validates and sanitizes incoming prompts, and consider adding explicit authorization checks before an agent forwards instructions. Monitoring MCP traffic for anomalous patterns and incorporating prompt‑validation logic at every hop can reduce the risk of cascading attacks. Treat the trust relationship inherent in MCP as a design decision that now requires security review.

