Production Learning Review

Workflow Loop Detection in Multi-Step Agent Runs

Detecting loops requires watching trajectory patterns, not individual steps.

Senior Staff Writer · · 11 min read
Cover illustration for “Workflow Loop Detection in Multi-Step Agent Runs”
Behavioral Drift · October 5, 2026 · 11 min read · 2,510 words

Workflow loops in multi-step agent runs are a distinct failure class from single-step errors that requires a different kind of detection. A loop is defined by what stays the same across steps, not by what breaks in any one step, which makes it invisible to the monitoring most teams already have in place.

Workflow loops as a distinct failure class

A single-turn error is a bad output produced at one step in a process. A workflow loop is a trajectory that never reaches a new state, despite continuing to execute for many steps afterward. The agent stays syntactically active the entire time: spans appear in the trace, tokens accumulate, tool calls fire on schedule. Semantically, though, nothing advances. The pattern visible in traces when this happens is consistent: identical or near-identical agent.trajectory.step spans, a token cost that keeps climbing, the same tool names called over and over, identical argument hashes attached to those calls, the same assistant intent restated turn after turn, goal progress that stays flat, and no terminal state reached.trajectory.step` spans, a token cost that keeps climbing, the same tool names called over and over, identical argument hashes attached to those calls, the same assistant intent restated turn after turn, goal progress that stays flat, and no terminal state reached.

Monitoring built around the final answer misses this. A final answer can look perfectly reasonable even when the trajectory that produced it burned dozens of unnecessary steps getting there, and a system that only checks outcomes has no way to see the waste embedded in the path. The failure compounds as systems grow more complex, because a single request in a modern agent deployment can cross a planner, a retriever, a tool executor, a memory store, and a handoff policy on its way to completion. A loop can cross every one of those layer boundaries and still look active, even correct, at each boundary you check in isolation. One agent retries a failing tool call. It hands off to a second agent carrying the same incomplete state forward. That second agent, finding nothing new to work with, hands the task back unchanged. Each individual handoff passes whatever local check would normally confirm that it worked. The loop lives in the sequence, not in any single link of it, and that is the property that makes it a different class of failure from a wrong answer at one step.

Where loops originate in the harness, not the model

Most loop triggers trace back to two layers of the harness surrounding the model, not to the model's reasoning itself. The first is the planner prompt. A planner instructed to "keep trying" without a step budget, and without a stop condition that recognizes repeated failure as failure, will faithfully replay the same action as many times as it's allowed to run. The agent asks a tool for missing data, gets back an empty or ambiguous response, interprets that response as confirmation that the data is still needed, and calls the same tool again with the same arguments. The planner is doing what its prompt told it to do, and that instruction is the problem.

The second major source is tool schema drift. When the schema an agent was given at registration diverges from the schema the server actually enforces, the agent doesn't receive an error it can reason about. It gets a silent mismatch or a malformed request that, from the outside, looks like a model problem rather than an infrastructure one, and the agent frequently retries rather than surfacing the failure, even though retrying produces the identical malformed parameters and therefore the identical failure every time. Drift of this kind can be effectively invisible to standard checks: the description text attached to a tool stays the same word for word while the underlying parameter fingerprint changes underneath it, so a schema validator scanning for textual differences reports nothing wrong. The downstream effect is that the agent falls back to a secondary tool in an attempt to "find" data it technically already has access to, adding latency and noise to the trace without ever producing a crash that would flag the real issue.

Beyond these two, handoff policies that return an unresolved task to the originating agent with the same incomplete state create loops, and every individual agent in the chain looks correct because the defect lives in the handoff logic connecting them, not in either agent's own behavior. Retry configurations that treat every tool error as transient, rather than distinguishing a temporary network blip from a persistent schema mismatch, will retry a failing call indefinitely instead of escalating it. If memory stores fail to update between iterations, the planner re-enters the identical decision branch on each pass, because from the planner's vantage point nothing has changed since the last step. The attribution matters in a practical sense: a model failure might point a team toward retraining or prompt redesign at the reasoning level, while a harness failure can usually be fixed in a prompt, a schema definition, or a retry config without touching model weights. Research on multi-tool orchestration in LLM agents backs this distinction, showing that long-horizon agent failures cluster in planning, execution feedback, and environment interaction rather than in the model's core reasoning.

Why agents cannot reliably self-report that they are looping

Agent frameworks that delegate the continue-or-stop decision to the model's own progress reports are building that decision on an unreliable signal. Research evaluating progress reporting across deployed models, using the public τ²-bench benchmark alongside a controlled testbed called StageIF with checkpoints placed across a task's lifecycle, finds that reporting reliability depends heavily on which stage a task has reached. Almost every model tested is reliable at some stages of a task and unreliable at others, and where the breakdown happens differs by model generation. Most deployed models lose accuracy specifically once work is actively under way, which happens to be the exact phase where a loop is most likely to form and go unnoticed. The newest generation of models closes that mid-task accuracy gap, but trades it for a different problem: these models grow conservative at the point of task completion, which is its own kind of reporting error and one that can mask a loop just as effectively.

The concrete failure plays out in two directions. A model can declare completion while the actual goal remains unmet, and if a runtime trusts that declaration, it will stop prematurely. Or a model can report "still working" through many repeated, nearly identical steps, with no external signal ever stepping in to interrupt it. A documented case from a telecom task on τ²-bench illustrates the first failure mode concretely: the user simulator sent a ###TRANSFER### token before the agent had actually performed the transfer, and the runtime stopped on that signal with the underlying goal unmet. The implication for anyone building these systems is that stop conditions belong to the runtime, grounded in observable trajectory state, things like step count, argument hashes, tool name repetition, and goal progress scores, rather than delegated to whatever the model says about its own status.

The trajectory-level signals that make a loop detectable

Detecting a loop requires comparing state across steps at the trajectory level. Instrumenting individual spans carefully is necessary groundwork, but it isn't sufficient on its own, because the diagnostic signal lives in the pattern that emerges across many spans rather than in the content of any single one. The pattern that logs need to surface includes a high iteration count for the task type, the same tool.name appearing repeatedly, identical argument hashes attached to those repeated calls, the same assistant intent restated across turns, goal progress that has gone flat, and a timeout rate that climbs as the run continues.

Two evaluator-level scores sit alongside raw trace inspection, and they sharpen the picture considerably. StepEfficiency scores how many iterations an agent actually consumed against the minimum number a task should have required, and a sudden drop in StepEfficiency across a release cohort functions as a loop signal even in cases where the run eventually completes successfully. GoalProgress scores forward movement between iterations directly, and when GoalProgress flattens after, say, the third step while the same tool arguments keep repeating, the exact boundary where the loop began becomes locatable in the trace rather than just suspected.

At the dashboard level, several signals deserve ongoing tracking, with thresholds set per workflow type since a search agent, a refund-processing agent, and a coding agent each have a different normal step count and a different baseline cost profile. Iterations-per-trace, tracked as a p95 or p99 histogram rather than an average, reveals the long tail where loops concentrate. Token cost per trace, watched alongside task completion rate, reveals loops through a specific combination: cost rises while completion stays flat. Timeout rate and human-escalation rate work as proxies for user frustration, and they round out the picture, since both tend to climb when a workflow is quietly failing to terminate even though nothing has technically crashed.

Mapping loop signatures to the harness layer responsible

The same surface pattern in a trace, a repeated tool call, a repeated intent, a handoff that bounces back and forth, can point to a different harness layer depending on what else accompanies it, and misreading which layer is responsible leads directly to a fix that looks reasonable but doesn't hold once deployed. A signature where the same tool is called with the same arguments and returns a null result each time usually points to missing data or a missing record, and the correct fix sits at the tool's data contract or at the planner's stop condition: capping retries at some small number like two, requiring a new identifier before trying again, or escalating the task to a human rather than retrying indefinitely. A signature where the same tool, same arguments, produce the same error rather than a null result points somewhere different, toward a tool that's down or a schema mismatch, and the fix belongs in the tool schema itself or in the retry policy, adding a tool-failure cap paired with retry-with-jitter rather than allowing unlimited replay of a call that will never succeed as written.

A signature where the assistant's intent stays the same across turns but the tool arguments keep shifting slightly points to a stuck reasoning chain inside the planner prompt, and the fix there is a re-prompt or a few-shot example that explicitly models what the exit condition looks like. A signature where execution bounces between two different tools suggests the planner can't resolve which tool actually owns the action in question, and the fix is tightening the tool descriptions themselves or adding a few-shot example to the planner that disambiguates the two. A ping-pong pattern in agent handoffs, where a task travels between two agents without resolution, points to unclear handoff ownership, addressed with a handoff-depth cap combined with a dedicated resolution agent or an escalation path. A long retry sequence running up against a rate limit points to misconfigured backoff, and that fix belongs in the gateway or the retry configuration.

The specificity here carries real weight, because each of these fixes targets a different artifact: a prompt, a schema, a retry config, a handoff policy. Shipping the wrong one leaves the underlying loop fully intact while adding a layer of complexity on top of it. Outcome-level signals, the kind that just report a run as failed, tell a team that something went wrong without telling them where. Trajectory-level analysis is what turns that outcome signal into an actual diagnosis, one that points at a specific artifact a team can go change.

Misattribution of loop root causes by single-pass LLM judges in long traces

A common instinct, once a loop is suspected, is to hand the entire trace to an LLM and ask it to find the root cause. That approach fails specifically for loops, because a long trace contains many plausible loop explanations scattered across planning spans, tool spans, and observation spans, and if you read it in a single pass, you tend to settle on the first coherent story you encounter and stop looking for competing explanations further on. Research on structured agent-failure diagnosis describes single-pass raw-trace diagnosis as inherently unreliable, tracing that unreliability to LLMs' documented struggles with long contexts, positional sensitivity within a document, and the effects of context compression, rather than premature commitment being the whole story. A separate study on root-cause localization corroborates the same shallow-diagnosis pattern from a different angle, framing it as an inherent tendency to favor early plausible explanations over deeper causal exploration, a related finding though not identical to committing before finishing a full read of the trace.

The same trace, queried twice, can produce two different confident-sounding diagnoses from the same judge. If an attribution method isn't stable under repeated queries against identical input, it isn't reliable enough to drive a fix, because a team acting on it risks changing a prompt or a retry config based on a diagnosis that was essentially a coin flip dressed up as analysis. Wrong attribution at this stage doesn't just waste effort. It leaves the real loop in place, and a team ships a change aimed at the wrong artifact.

Validating loop fixes against historical traces before shipping

A harness change that eliminates a loop in local testing but hasn't been checked against the historical traces where that loop originally appeared is a hypothesis, not a fix, and shipping a hypothesis without replay validation is shipping blind to whether the actual production conditions that caused the loop have been addressed. The validation workflow starts by pulling the specific traces where the loop signature showed up, identified through argument hash repetition, a StepEfficiency drop, or GoalProgress flattening, and treating that set as a dedicated regression set rather than folding it into general evaluation. The candidate fix, whether that's an updated prompt, a corrected schema, a revised retry config, or a tightened handoff policy, then runs against that same regression set, with the bar for passing set at two conditions: StepEfficiency has to improve or at minimum hold steady, and loop detection has to stop firing on those specific traces. Any new deploy that touches the planner, the tool schema, or the handoff policy involved should gate on that regression set passing, rather than relying only on a synthetic golden set built ahead of time.

Shadow evaluation against live traffic catches regressions that a golden set built in advance will miss, because production traces expose argument combinations and state transitions that synthetic test inputs just don't generate. This matters especially for loop fixes, since loops tend to occur in exactly the edge-case state combinations that synthetic test design tends to omit by construction. Historical loop traces should be kept as a dedicated regression layer, separate from general quality evaluation, because loop failures carry a different signature from quality failures and need their own evaluator scores, StepEfficiency and GoalProgress specifically, along with dashboard signals like loop-rate broken out by workflow, to catch a regression before it reaches production again.

Sources

  1. The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
  2. The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?
  3. When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents
Filed underBehavioral Drift

More in Behavioral Drift