The Agentic Harness

Prompt Injection Detection in Agentic Pipelines

Prompt injection succeeds because agents can't separate instructions from retrieved data.

Staff Writer · · 13 min read
Cover illustration for “Prompt Injection Detection in Agentic Pipelines”
Agent Failure Diagnosis · September 30, 2026 · 13 min read · 2,992 words

Prompt Injection Detection in Agentic Pipelines.

Prompt injection in agentic pipelines: a structural problem, not an input-filtering one

Prompt injection in an agentic system is not a bug you patch at the front door. Engineers who treat this as a single filtering problem at the point where the user types a message will catch the easy cases and miss almost everything else.

The comparison to SQL injection gets made often, and it's useful precisely because it shows where the analogy breaks down. Both exploits share a root cause: the system fails to separate trusted instructions from untrusted data. SQL injection got a structural fix because a database can enforce a syntactic boundary between a query and its parameters, parameterized queries close the hole permanently, not just probabilistically. Nothing like that exists at the model layer. A system prompt, a user's question, and a paragraph pulled from a retrieved document all arrive at the model as the same thing: natural-language tokens. There's no delimiter the model is architecturally required to respect, no equivalent of a prepared statement that keeps instruction and data in separate channels. The model reads everything as language, and language is exactly the medium an attacker needs to hide a command inside.

That's why OWASP lists prompt injection as LLM01:2025, the top-ranked risk in its Top 10 for LLM Applications. A May 2026 industry survey from Zylos AI put the figure at 73% of production AI deployments carrying this exposure, a number that matters for scale even though it comes from a vendor survey rather than from OWASP itself. Whatever the exact prevalence, the direction of the finding lines up with what security researchers keep observing: this isn't a rare edge case, it's close to the default condition of shipping an agent.

Simon Willison's "lethal trifecta" framing captures the structural risk cleanly. The trouble is that most production agents already have all three. A customer support agent wired into a CRM, reading user-submitted text, and able to send email is the textbook case, and it's also just an ordinary support bot. Nobody built that agent recklessly. It's simply what a useful support agent looks like, which is exactly the point: the vulnerability isn't a design mistake, it's a consequence of giving the agent the capabilities it needs to do its job.

This reframes what engineers should actually be solving for. Instead of asking how to prevent injection outright, a better use of effort is asking what limits the damage once injection succeeds, because for most production agents the exposure is baked into the architecture, not an artifact of sloppy coding. The 2026 edition of the OWASP GenAI Security Project's State of Agentic AI Security and Governance report reflects this shift directly: it catalogs CVEs, vendor advisories, and production incidents across nearly every category of agentic risk, a marked change from the 2025 edition, which still discussed these threats as hypothetical. The theoretical period is over. What follows is a layer-by-layer account of where injection actually enters agentic pipelines, and why the fix that works at one layer routinely fails at the next.

The evolution of the threat: from chatbot jailbreaks to goal hijacking in multi-agent systems

Direct injection is the oldest and simplest form: an attacker with access to the model's input field types something like "ignore all previous instructions" and hopes the override sticks, or attempts role-play hijacking to coax the system into a jailbroken persona. For a standalone chatbot with no tool access, a successful direct injection is embarrassing. For an agent with tool access, the same attack produces real-world consequences: a sent email, a deleted file, a wire transfer.

Indirect injection is the more consequential development, because the payload doesn't come from the user. It arrives embedded in content the agent retrieves as part of doing its job: a PDF, a webpage, a database record, an MCP tool description. Once an instruction and a piece of retrieved data are both flattened into the same token stream, the model treats them identically. Google's own monitoring of the web found a 32% rise in malicious prompt injection payloads embedded in web content between November 2025 and February 2026, showing this is a live, growing category, not a theoretical curiosity confined to red-team papers.

Goal hijacking sits a level above both. Rather than triggering one bad action, the attacker redirects an autonomous agent's entire objective across a multi-step task. In a multi-agent pipeline, that hijack doesn't stay contained. A compromised agent can instruct downstream agents, poison shared memory, or manipulate an orchestrator's decisions, and the resulting damage scales with how much autonomy and tool depth the pipeline grants. OWASP's Top 10 for Agentic Applications maps prompt injection onto six of its ten risk categories, the clearest evidence that it operates as a family of distinct vulnerabilities scattered across the framework.

The consequences aren't abstract. In June 2025, a banking agent lost $250,000 to an attack that used zero-opacity white text embedded in a customer support chat to bypass transaction verification, a documented production failure, not a lab demonstration. The supply chain has its own version of this story. In March 2026, a backdoored version of the LiteLLM package sat on PyPI for roughly 46 minutes and still racked up nearly 47,000 downloads, pulled in as a transitive dependency by CrewAI, DSPy, Microsoft GraphRAG, and other frameworks. The attack surface for an agentic pipeline now includes every package the harness depends on, whether or not anyone on the engineering team ever reviewed that dependency directly.

Layer one: detecting injection at the user input boundary

Direct injection at the input boundary is the layer the industry understands best, and it's the one where existing guardrail products perform reliably, because pattern analysis and semantic classification are genuinely effective against explicit override attempts, jailbreak phrasing, and role-manipulation prompts. Two provider approaches illustrate this concretely. Azure's Content Safety Prompt Shield runs a dedicated model trained to detect jailbreaks and indirect injection, evaluating the full message payload and returning an attackDetected flag separately for the user prompt and for each document in a documentsAnalysis field. AWS Bedrock Guardrails takes a broader approach, applying pattern-based and semantic checks at both the input and the output layers, which matters because, as later sections make clear, output-layer scanning is not optional.

Where this detection runs matters as much as how it works. Application-layer defenses, a regex check bolted onto a request handler, a moderation API call sprinkled into application code, tend to fall apart once a system grows past a handful of services. Every new microservice or model integration becomes a fresh place where the check might have been forgotten. Credentials for the moderation service sprawl across dozens of codebases. Audit logs fragment into inconsistent formats scattered across different teams' logging conventions. A gateway-layer control point resolves all three problems at once: every request funnels through a single process, credentials live in one place, and the audit trail looks the same no matter which workload generated it. Concretely, this means CEL-based rule targeting and tool allow-lists enforced at the gateway, which block injection-driven tool abuse without touching a single line of application code.

None of this, however well executed, reaches past the front door. Input-layer detection is structurally blind to anything that enters through retrieved content, tool output, or memory, because those channels never pass through the user input field. It's a consequence of where the checkpoint sits in the pipeline.

Layer two: injection through retrieved content and RAG pipelines

Indirect injection through retrieval works because the model has no mechanism for treating fetched content with more suspicion than a direct instruction. A document, a webpage, an email, a database record: whatever the agent pulls in as part of its normal task arrives as content, and the model processes it the same way it processes an instruction, because there is no syntactic marker distinguishing the two. Research from January 2026 found that five carefully crafted documents, planted anywhere in a retrieval corpus, were enough to manipulate AI responses 90% of the time through RAG poisoning, a trivially achievable bar for a motivated attacker. That's not a low-probability tail risk worth a footnote.

The documented cases give the abstraction some teeth. A Google Docs file was used to trigger an AI IDE agent into fetching and running a malicious Python payload hosted on a GitHub Gist, harvesting secrets with no user interaction required at all. Payloads have embedded complete PayPal transaction specifications inside ordinary web content, aimed at agents with payment capabilities. Injection text has been hidden in website meta tags specifically to route AI-mediated financial actions toward attacker-controlled Stripe donation links. Code repository comments have been crafted to push coding assistants with shell access toward destructive file operations. And GitHub Copilot's CVE-2025-53773, rated 9.6 on the CVSS scale, showed the mechanism at its most severe: malicious instructions embedded in externally fetched content led the agent to execute attacker-controlled commands, a full remote code execution chain built entirely out of content the agent was simply asked to read.

Detecting this requires a different posture than input filtering, because the injected content arrives labeled as trusted retrieved data, there's no prior signal telling the system to be suspicious of it. Content mutation detection catches responses that have been subtly steered off course by injected instructions, output that reads as plausible on its face but has quietly departed from what the task actually called for, a category no pattern-matcher would ever flag. Decision trace and lineage tracking give investigators something to work backward from: linking every agent output to the specific upstream assets that produced it means that when a poisoned context entry does influence a decision, the lineage trail shows exactly where it got in. Taint tracking closes the loop by marking data from untrusted sources and following it through the reasoning chain. If tainted data is about to drive a high-privilege action, the system can demand confirmation or block the action outright. What unites all three approaches is that they operate on behavior and provenance rather than on the content of the input itself, because the input, by the time it reaches this layer, already looks completely legitimate.

Layer three: injection through tool outputs and the emerging tool-integration surface

MCP widened the attack surface considerably, because it opened up new channels through which injected content reaches the model without passing through anything resembling a user input field. Tool descriptions are one channel: the metadata a server exposes, including tool names and description fields, gets read by the model during discovery, and an attacker who controls that metadata controls a piece of the model's context before a single tool call has even happened. Tool output is the second channel, and arguably the more dangerous one: a compromised or malicious MCP tool can return adversarial content in its response, and that content enters the model's context carrying the implicit trust of legitimate tool output.

Two CVEs illustrate how this plays out in practice. CVE-2026-22708, against Cursor, let an attacker poison the agent's execution environment so that allowlisted commands like git branch delivered arbitrary payloads, the exploit worked by abusing shell built-ins such as export to bypass the allowlist entirely Prompt injection still drives most agentic AI security failures in production - Help Net Security. CVE-2025-59532, against OpenAI's Codex CLI, let the agent's own output redefine the boundary of its sandbox, a strange inversion where the thing being protected against becomes the thing doing the protecting Prompt injection still drives most agentic AI security failures in production - Help Net Security.

The supply chain for MCP servers has its own cautionary tale. A package called postmark-mcp shipped fifteen clean versions, building a track record of apparent legitimacy, before quietly inserting a single line of exfiltration code, the first documented case of a malicious MCP server caught in the wild. CVE-2025-6514, rated 9.6 on CVSS, was a remote code execution flaw in mcp-remote, an npm proxy package with over 437,000 downloads, a reminder that popularity is not evidence of safety.

Detection here leans heavily on goal alignment verification: before a tool call executes, check whether the action plausibly serves the user's stated goal. An agent asked to summarize a report has no legitimate reason to email that summary to an external address, and that mismatch is often enough to flag a large share of successful injections. Microsoft's 2026 defense architecture builds this into infrastructure directly, treating every tool invocation as a high-value, high-risk event: before execution, context gets sent via webhook to Microsoft Defender, which analyzes intent and destination and returns an allow or block decision in real time. MCP tool allow-lists remain a useful preventive layer on top of this, explicitly enumerating which tools an agent may call so that even a successful injection inside the model can't reach a tool that was never authorized. But allow-listing which tools can be called is a different guarantee from trusting what those tools return: a tool being on the allowlist says nothing about the integrity of its output on any given invocation.

Research from ETH Zurich put numbers to how exploitable this layer already is Assessing Automated Prompt Injection Attacks in Agentic Environments. Testing across 80 task pairs spanning workspace, banking, travel, and Slack domains, the black-box TAP method achieved a 45.2% attack success rate against Qwen3-4B, compared with 24.1% for the gradient-based GCG method Assessing Automated Prompt Injection Attacks in Agentic Environments. GPT-5 held up far better, but even there TAP achieved roughly a 5% attack success rate and 30% under a Success@N metric Assessing Automated Prompt Injection Attacks in Agentic Environments. Attackers targeting tool-calling pipelines already have automated methods that work at meaningful rates against production-grade models, not just against toy benchmarks.

Layer four: injection through persistent memory and cross-session poisoning

An agent that remembers prior conversations can have that memory poisoned in one session and carry the corruption into every session that follows. This is qualitatively different from the earlier layers. A poisoned memory entry does damage on a delay, activating whenever some future session happens to read it.

In multi-agent pipelines, this compounds further. A hijacked agent can write a poisoned entry into shared memory that later agents read as ordinary state, and downstream orchestrator decisions end up shaped by that poison without any new user interaction ever occurring. The attack happened earlier, and the system carried it forward faithfully, exactly as memory is supposed to work.

Detection strategies from the earlier layers mostly don't transfer here, and here's why. Input-layer classifiers are useless because the memory content was written by the agent itself, it arrives already labeled as trusted internal state, not as external input that would draw scrutiny. Content mutation detection struggles too, since the injected instruction may have been written in a prior session and looks, to any check happening now, exactly like legitimate accumulated memory. Goal alignment verification runs into a timing problem: a memory write made in good alignment with the user's goal at the moment it was written is indistinguishable from a poisoned one, if the attacker chose the moment of injection carefully.

What does transfer, in adapted form, is lineage tracking applied to memory writes specifically, giving every memory entry a provenance record: which session wrote it, from what source content, at what trust level, so that an audit trail exists if a memory-influenced decision later gets flagged. Taint propagation extends the same way: a memory write influenced by tainted retrieved content inherits that taint, and later reads of that entry trigger confirmation requirements before it's allowed to drive a high-privilege action. Memory consistency checks add a third angle, scanning stored memory for patterns that conflict with or quietly override system-level instructions, since injection that rewrites behavioral constraints tends to leave a semantic fingerprint that's detectable if anyone's looking for it. What separates this layer from the others is that per-request scrutiny isn't enough on its own. Memory needs a standing audit process running against the store itself, independent of whatever happens to be in front of the model on any given call.

Output scanning as a required layer, not an optional addition

Prompt injection in agentic systems is very often aimed at getting data out, not just at getting a bad action performed. The injected instruction causes the model to fold sensitive information into its output, and that output then travels somewhere, to a user, to a downstream agent, to an external API call. Blocking only at the input layers misses this pattern completely, because an injection that entered through a RAG document can produce reasoning that looks entirely reasonable right up until the final output quietly contains the exfiltrated data.

The lethal trifecta framing makes the case for output scanning almost by definition. Any agent holding access to private data, exposure to untrusted content, and some path for data to leave the system will, sooner or later, route exfiltrated content through its own output. Scanning that output is the last point in the pipeline where the data can still be stopped before it leaves. If it is missed there, no further checkpoint remains to catch it.

Output scanning catches categories nothing upstream can. Content mutation, output that reads as plausible but has been quietly steered away from the intended task, never touches an input classifier, because the manipulation happened entirely inside the model's own reasoning. Data exfiltration payloads, sensitive tokens folded into otherwise ordinary-looking text, only become visible once the output actually exists. In multi-agent systems, a poisoned output from one agent becomes the poisoned input for the next. Output scanning at each stage is also, functionally, input protection for whatever comes after it. Treating the four layers, input, retrieval, tool output, and memory, as separate problems each solved by their own detection method is not an academic distinction Assessing Automated Prompt Injection Attacks in Agentic Environments. A security posture that actually holds looks very different from one that only looks like it does until the first document nobody thought to check.

Sources

  1. Prompt injection still drives most agentic AI security failures in production - Help Net Security
  2. Assessing Automated Prompt Injection Attacks in Agentic Environments
  3. Prompt Injection at Scale: Defending Agentic Pipelines Against Hostile Content - TianPan.co

More in Agent Failure Diagnosis