Detecting Prompt Injection in Production LLM Agent Pipelines
Securing agents requires monitoring every data entry point, not just user input.

Prompt injection in production agent pipelines is a problem of every point in the pipeline where outside content enters the model's context window, and detecting it requires instrumenting the full harness, not just the front door.
Prompt injection as an agent-layer problem
Large language models process every token in one stream, at the same layer, with no runtime mechanism that separates a system instruction from a sentence lifted out of a customer email. Vectra's analysis of the LLM processing pipeline describes this: by stage four, the system prompt, the user's input, and any external context the model has pulled in all arrive with equal weight. The model has no built-in way to mark one as privileged and the other as data to be read, not obeyed.
That structure should sound familiar to anyone who has worked in application security. It is the same failure mode behind SQL injection: code and data mixed in one channel, with the interpreter unable to tell which is which. SQL injection and this version of the problem differ in scope. SQL injection lived in the query layer of a database application. This version of the problem touches every AI application that takes in outside input, because outside input is, structurally, how these systems work.
The stakes grow with the pipeline's reach. Picture an agent that reads a Jira ticket, pulls a Confluence page for context, queries a database, and posts a summary to Slack, all inside one user turn. Each of those four steps is a place where text enters the model's context from a source the user never typed into and a gateway watching the user's prompt never sees. A guardrail sitting at the chat input catches nothing happening three steps downstream, because by the time the Confluence page is fetched, the attack surface has already moved past it.
That is why security researchers split prompt injection into two distinct classes, each with its own defenses. Direct injection arrives in the user's own prompt, and pattern-based and semantic classifiers can catch a meaningful share of it because the attacker's text sits in a channel someone is actually watching. Indirect injection arrives in content the agent retrieves on its own, a ticket, a page, a row in a database, and the attacker never has to touch the user-facing channel. The instructions ride in on data the system already trusts. If the injection surface runs through every layer of the harness, the only way to catch it is to watch every layer, not just the one the user types into. That is the argument the rest of this piece makes.
The full injection surface an agent pipeline exposes
Mapping that surface starts with the protocols agents now use to call tools and retrieve context. The Model Context Protocol and similar function-calling harnesses have widened the indirect injection surface to at least four distinct entry points, and each one moves through the pipeline differently.
Maxim AI's guide to this problem lays the four out clearly. Tool descriptions are the first: MCP server metadata can carry instructions buried in a tool's name or its description field, and the model reads that metadata during tool discovery, before a user has said anything. Tool output is the second: a compromised MCP tool can hand back adversarial content that the model treats as trustworthy, factual data simply because it came from a tool call rather than a stranger. Memory stores are the third: an agent's persistent memory can be poisoned once and then quietly shape every session that follows. RAG retrieval results are the fourth: any document sitting in the retrieval corpus, a wiki page, a support ticket, a PDF someone uploaded months ago, can carry override instructions that activate the moment it gets pulled into context.
What makes these four entry points hard to defend uniformly is that each one carries its own trust assumption, baked into how the harness is built. Tool descriptions get read before any user turn even starts, so there is no user-channel checkpoint to intercept them. Tool output is treated as ground truth the moment it returns. Memory is treated as the agent's own accumulated state, not as external input. RAG content is treated as authoritative, retrieved fact, because surfacing trusted information was why it was fetched. A defense built for one of these assumptions does not transfer cleanly to the other three.
Longer context windows make injected text harder to detect. LongPIBench, a benchmark covering four realistic scenarios, paper peer review, resume screening, code review, and email summary, tests injection defenses across context lengths running from a few thousand tokens into the tens of thousands. The results show that defenses performing well on short-context benchmarks degrade substantially once the context grows to the length enterprise agents actually work with. Even simple heuristic attacks post high success rates in these long-context settings when no defense is present, and named defenses including MetaSecAlign lose effectiveness as the context window grows. A production agent reading a full document corpus or a lengthy email thread is operating in a setting most published defense evaluations never tested.
Purpose-built agents carry a further complication. These systems enforce their own application-level rules on top of whatever general safety training the base model has, rules like rejecting unauthorized queries or restricting which records a user can touch. Research on PVDetector points out that prompt injection can defeat these purpose-specific restrictions even while the model's general safety alignment holds, because the attack targets the application policy layer sitting above it, a layer that general-purpose guardrails were never built to watch.
Why gateway-level guardrails alone cannot close this surface
A gateway or input guardrail does real work. It is one layer in a stack that needs several, and it is not built to see the surfaces just described. The limitation is structural, a matter of how the system is built rather than a given vendor's filters being weak.
Maxim AI's guide names three specific conditions where application-layer defenses break down at scale. The first is coverage gaps: every new microservice or agent added to a system has to independently implement its own injection checks, and in practice, some will and some won't. The second is credential sprawl: the access needed to run guardrail checks has to be managed separately across every service that implements them, multiplying the places a misconfiguration can live. The third is fragmented audit evidence: without a unified log, there is no single place to see which requests triggered a violation, which user submitted them, or which downstream application got hit. A successful indirect injection can produce output that scores perfectly well on ordinary task-completion metrics while quietly exfiltrating data or triggering a tool call nobody asked for. The monitors watching output quality have nothing to flag.
OWASP's own guidance on LLM01:2025 states the underlying reason directly: given the stochastic influence at the heart of how these models work, it is unclear whether any fool-proof method of preventing prompt injection exists. Maxim AI's guide draws the same conclusion from that same source: no single technique eliminates injection, and the right architecture stacks multiple independent layers, each one raising the cost of a successful attack rather than promising to stop it. A major model provider reached the same place from its own side of the problem. OpenAI's Lockdown Mode for ChatGPT, launched February 13, 2026, came with a public acknowledgment that prompt injection in AI browsers "may never be fully patched."
The risk compounds once an agent is wired up with tools it can call on its own. Vectra's analysis calls this "agentic amplification" and traces what researchers now describe as a promptware kill chain, borrowing the structure of the traditional cyber kill chain to describe prompt injection as a multi-stage attack mechanism rather than a one-shot trick. A single injected instruction, once it lands, can set off a chain: exfiltrate data, execute code, move laterally to another system. Once that chain starts moving, no single checkpoint downstream is positioned to stop all of it.
Two disclosed incidents show the chain completing in practice. In August 2024, researchers at PromptArmor demonstrated a Slack AI exfiltration attack in which injected messages caused the summarizer to pull private channel content into a response visible to an attacker. Between 2024 and 2026, researchers also documented injection attacks that steered coding agents, including GitHub Copilot and Cursor, into actions the user never requested. In both cases, the attack ran its course before any single detection point along the way had a chance to catch it.
What production traces reveal that no upstream guardrail can
If gateways cannot see past the user channel, the question becomes what can see the rest of the pipeline after the fact. The trace answers it: a record of every prompt, every guardrail verdict, every tool call, and every retrieved chunk, captured as the agent actually executes. A successful injection leaves a trail in that record, and the trail is what attribution runs on.
Span-based trace replay has become a basic requirement for watching agents in production, for a simple reason: ordinary logs cannot reproduce a failure that was non-deterministic to begin with. A trace that captures the arguments passed to a tool, the raw output that tool returned, and the model's next decision shows exactly which layer carried the instruction that shouldn't have been there. That is a different kind of evidence than a log line saying a request failed. It shows the sequence: what the model was told, what it decided to do next, and where in that sequence something diverged from what should have happened.
OpenTelemetry's GenAI export paths make this practical across different systems rather than locking it to one vendor's format. The trace format captures an agent's full loop, its model calls, its tool calls, and its retrieval operations in one unified record. Infrastructure-level spans, HTTP requests, database queries, and the rest, can be parented into that same trace using standard OpenTelemetry instrumentation already common in backend engineering. That connection matters because it lets an engineer tell a model-quality failure apart from an application bug or an infrastructure fault, three categories of problem that look identical from the outside but require entirely different fixes.
Attribution gets harder, not easier, as these systems scale up. As execution logs grow larger and more distributed, models doing automated root-cause analysis tend to settle on a plausible-sounding explanation before they have worked through the evidence. Hierarchical causal graph methods address this directly. CHIEF, for instance, builds a causal graph of the execution and uses oracle-guided backtracking to prune down the space of possible causes, then applies counterfactual attribution separately to tell a genuine root cause apart from a symptom that merely propagated from somewhere else. The method forces a systematic walk through the evidence instead of letting the analysis stop at the first answer that looks reasonable.
Instrumenting each harness layer for injection-aware tracing
Turning that argument into practice means placing instrumentation at the exact point where each surface contributes tokens to the context window, not only at the point where the user's message comes in. The work breaks down by harness layer: prompt, tool, workflow, and memory.
On the tool side, every tool's name, description, and schema gets stamped into the prompt on every request you send. That makes tool schema drift, an upstream API quietly changing a parameter name or a required field without a corresponding update to the harness, a silent enabler of injection: the model may route a call incorrectly on its own, or injected content may steer it into calling a tool with argument shapes nobody intended. Tracing needs to capture the full schema at the moment of the call, not just the tool's name, so a drift event is visible after the fact.
Tool output needs its own discipline. The span recording a tool's response has to capture the raw content before any post-processing touches it, because a compromised MCP tool can hand back adversarial content that the harness treats as trustworthy data the instant it arrives. The injection event lives in that response rather than in the request that triggered it, so a trace that only logs the request misses it.
RAG retrieval needs each chunk treated as a traceable object in its own right: its content, its source document, and its position in the context window the model actually saw. Without that, a behavioral change downstream has no way to be traced back to the specific document that caused it.
Memory reads need the same treatment. Agents with persistent memory treat prior session state as already trusted, so a poisoned memory store can shape future sessions without any user-turn injection. Tracing memory reads as their own spans is what makes that propagation path visible instead of invisible.
MCP tool description discovery deserves particular attention because it happens before any user turn begins. Maxim AI's guide flags MCP server metadata, the tool name and description fields the model reads during discovery, as an injection surface in its own right, one that standard setups frequently leave untraced.
Workflow loops are their own category of risk. Some coding agents can be forced back into a planning loop when a hook intercepts the model's attempt to exit and reinjects the original prompt into a fresh context window, and attackers can craft indirect injections designed specifically to trigger that re-entry. A span recording the loop's iteration count and the condition that triggered each re-entry makes runaway cost or latency visible in the trace before it turns into a runaway execution nobody notices until the bill arrives.
Determinism is worth measuring across replayed traces for the same reason. Even at zero temperature, the order in which streamed events arrive, the timing of tool responses, and small variations in tool output can produce different final answers from one run to the next. If an agent shows low determinism across replays, it is usually running on underspecified prompts or loosely defined tool descriptions, and both conditions make injection easier to pull off.
Detection methods that operate on what the trace reveals
No single detection method covers this entire surface. Effective detection stacks several independent techniques, each one doing its best work at the specific layer where it has the clearest signal.
Input-layer classifiers remain useful for direct injection. Pattern-based and semantic classifiers sitting at the user-input boundary catch known phrasings of direct injection reliably, so they belong in the stack even though they cannot see anything past the point where the user's message ends. Their coverage stops at the user channel, by design, and everything this piece has laid out about tool output, retrieval, memory, and tool metadata lies past that boundary. Catching what happens there requires the trace-level instrumentation described above, applied consistently across every layer of the harness, because the alternative is a detection system that only ever sees where the attack started and never sees where it actually did its damage.
Sources
- LongPIBench: A Long-Context Benchmark for Prompt Injection
- PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
- Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
- Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection


