Loop Latency and Its Effect on Agent Improvement Velocity
Sequential LLM calls create hidden delays that compound into agent improvement bottlenecks.

Every deployed agent gets better, or doesn't, through the same four-step sequence: observe a failure, figure out what caused it, change something in the system, and validate that the change actually worked before shipping it. That loop sounds simple enough on paper. In practice, each joint in that chain carries its own delay, and those delays don't behave the way they do in ordinary software development.
Classical software iteration has a comparatively short path from bug to fix. A stack trace points at a line of code, the defect is deterministic, and a patch either resolves the exception or it doesn't. Agent systems break that assumption at the root. Because model generation and tool execution in most agent frameworks run sequentially rather than in parallel, a workflow that stacks up many LLM calls per task turns each per-token cost into a per-task cost and each per-call delay into part of the total wall-clock time for that task. That serialization means the loop itself is slow to execute, and a system that's slow to execute is also slow to observe at any meaningful volume. The rest of this piece works through what that means at every stage of the improvement cycle, and where the delay compounds into something closer to a ceiling than a nuisance.
Observation latency buildup before a team knows something is wrong
The first tax comes before diagnosis even starts. A failure has to happen enough times, surface somewhere a human or a system will notice it, and get distinguished from ordinary noise before anyone treats it as a signal to act on. None of that happens instantly, and none of it happens automatically either.
Most observability tooling built for conventional software wasn't built to see what actually breaks in agent systems. Standard APM tools can't detect tool-call failures, context truncation, or runaway loops, which are the failure modes that dominate real agent incidents, unless the instrumentation is built with agents specifically in mind. A team running without that instrumentation is missing more than some detail. It's watching a lagged, incomplete version of its own system.
Even when telemetry exists, teams frequently stop at the wrong layer of it. The signal that actually explains a failure lives in which tools got called, what those tools returned, which LLM calls fired, which spans failed, and how each of those steps fed into the final output. Watching only the final output means missing nearly the entire causal chain, and that blindness extends the time between something breaking and someone noticing.
Tool schema drift, documented as its own failure category in Memory-Induced Tool-Drift in LLM Agents (arXiv:2605.24941), produces no exception but corrupted outputs. No exception fires. Outputs just get worse, and they keep getting worse across production cycles until the degradation is large enough to register as a real signal. Context accumulation works the same way from a different angle: stale tokens pile up over a multi-turn session and crowd out what's actually relevant, degrading quality without tripping anything an observability hook would catch.
There's a third dimension here that's easy to miss because it looks like a feature rather than a bug. Research on the τ²-bench and StageIF testbed finds that nearly every deployed model is reliable at some stages of a task and unreliable at others, and most models (short of the newest generation) lose accuracy specifically in the middle of a task. Frameworks that lean on the model's own signals to decide whether to keep going or stop are, in effect, making loop-control decisions on data the model itself can't be trusted to report accurately. That makes it genuinely hard to tell a task that finished from a task the model just decided, incorrectly, was finished. The practical fix isn't philosophical: task flow shouldn't be governed by a model's self-reported state at all, which means observation needs independent runtime instrumentation instead of taking the model's word for where it stands.
Attributing a failure to the right layer as the hardest and slowest step in the cycle
Knowing something failed is close to useless without knowing which layer of the system caused it. Get that wrong, and the fix lands on the wrong component, the underlying problem persists, and the team is back at square one, running the diagnosis clock over again. A misattributed failure that triggers a harness rewrite when the actual cause was an ambiguous prompt is exactly how a two-day debugging session turns into a two-week engineering project.
Part of what makes attribution so slow is structural. Execution logs in agent systems are large, scattered across components, and written substantially in natural language rather than clean structured data. When those logs grow long, LLM-based judges tasked with analyzing them tend to lock onto the first plausible explanation rather than continuing to dig. The Continual Search framework was built specifically to counter that tendency, nudging the judge across multiple turns to keep examining evidence it hasn't resolved yet. Even a fully automated attribution tool needs deliberate architecture to avoid quitting early, which says something about how strong the pull toward premature conclusions really is.
Multi-agent systems raise the difficulty further. When multiple LLM-powered agents interact with each other, with external tools, and with their own internal reasoning all at once, the resulting logs tangle in ways that make isolating the decisive failure point structurally harder. Recent research has started formalizing responses to exactly this problem. AgentRx and Who&When Pro both aim to identify which agent and which specific error step caused a multi-agent failure, with Who&When Pro built on more than 12,000 labeled trajectories spanning different agent frameworks, domains, and modalities. A separate interaction-centric taxonomy frames the whole question around one distinction: was this a failure of the model, or a failure of the harness around it.
The two failure types look identical from outside the system, and that's what makes attribution hard. An agent stuck in an infinite loop because it genuinely can't recognize task completion, and an agent stuck in a loop because the harness has a bug in its exit condition, produce the same symptom. Fix the wrong one, and the entire improvement cycle restarts from zero. Large-scale study across hundreds of agent trajectories has started producing a structured way to sort this out, decomposing each run into operational modules, memory, reflection, planning, action, and system-level operations, and attributing failures to whichever module actually caused them. That work, published as the AgentErrorTaxonomy, is less an academic exercise than a direct answer to how much time gets wasted guessing.
The harness failure categories most frequently misdiagnosed
Some harness failures are genuinely hard to catch because they wear a model's clothing.
Schema drift is the clearest case. A downstream API changes the shape of its response, the tool parser misreads the new shape, and no exception gets thrown anywhere in the pipeline. Outputs just get quietly worse over time, accumulating across production cycles, and the instinct across most teams, understandable and almost always wrong, is to blame the model's reasoning before anyone thinks to check the tool contract. That instinct is understandable and almost always wrong.
Prompt ambiguity fails in a similar shape. The model produces something internally coherent, something that reads as a reasonable answer, but wrong for what the task actually needed. Without a clear reference for what the prompt was supposed to mean, that failure looks like the model hit a capability ceiling rather than the more mundane truth: the instruction wasn't specific enough. Teams in this position tend to reach for a model upgrade before they reconsider how the prompt was written, which is exactly backward when the prompt is the actual defect.
Workflow loops carry the highest cost of all when misdiagnosed. A loop that never terminates can come from a harness control-flow bug, where the loop condition is never satisfied, or from a model that cannot recognize task completion, a model progress-reporting failure per Wang et al.'s findings on the lost-mid-task pattern. Treating one as the other burns through an entire diagnosis-fix-validate cycle for nothing.
Context compaction rounds out the list, and it's arguably the sneakiest of the four because it erases constraints rather than corrupting outputs directly. By the time a multi-turn session has accumulated enough stale context to visibly degrade quality, the failure reads like the model drifting off-task rather than what it actually is, a memory management failure inside the harness. Sessions can grow to several times their original token count by the later turns, and most of that growth is stale material doing no useful work.
The throughline across all four is the same. Each one is a harness problem wearing a model problem's face, and every time a team treats it as the latter, the cost isn't just a wrong fix. It's a full wasted cycle through observation, attribution, and validation, repeated from scratch.
The harness and why changing it is slower than changing code
The harness is everything wrapped around the model that isn't the model itself: the prompts, the tool definitions, the control flow deciding what happens next, the memory and retrieval systems, the context management logic governing what stays and what gets dropped. Changing any single piece of that can send effects rippling through the rest in ways that aren't obvious from looking at the change in isolation, which makes the scoping of a harness fix slower and riskier than scoping an equivalent code fix.
The scope of what counts as "the harness" has also grown. As context windows filled up and prompt templates became unwieldy on their own, engineering effort shifted away from prompt tuning and toward the infrastructure sitting around the model: external memory stores, tool registries, protocol definitions, sandboxes, sub-agent orchestration layers, compression pipelines, and evaluators. Each of those is now its own change surface, with its own deployment cycle and its own way of breaking.
That interdependence is the real reason harness changes move slowly even once the correct fix has been identified. Adjusting a prompt shifts the model's output shape. That shift changes what the tool parser expects, which can quietly invalidate whatever gets written to memory downstream. None of that is visible from the diff alone, and that is why the scope of a harness change is so often larger than it looks at first glance.
Evals belong inside that same list of harness components, not bolted on afterward. Tools and sandboxes expand what an agent can do, but without evaluators wired directly into the harness there's no automated way to tell whether a given change actually helped or just added complexity. Teams without that in place end up paying a manual QA tax on every single change they make, and that cost is exactly what the next stage of the loop runs into.
Automated approaches to harness evolution, systems that propose edits to prompts, tools, or loop structure, test them against a benchmark, and keep whatever scores better, are an active area of research right now. The underlying insight to carry forward, though, is that automation reframes the problem rather than solves it. It's that harness changes need an evidence-backed, structured proposal-and-score cycle rather than one-off edits followed by someone eyeballing the output. Which raises the obvious next question: what does scoring a change well actually require before it ships?
Validating a harness change before shipping as the step teams most often skip or cut corners
Shipping a harness change without checking it against real production traces is shipping blind. A change that scores well on a synthetic benchmark, or on a small set of examples someone hand-picked, can regress quietly against the actual distribution of live traffic, and when it does, the team is back at the start of the loop, paying for a full cycle of observation and diagnosis it thought it had already finished.
The standard pre-ship gate for catching this is golden-set replay: a curated collection of traces with known expected outputs or rubric scores, run against the current pipeline on every deploy, with aggregate scores tracked over time. Work on Chronicle formalizes this specific pattern as cut-point replay for regression testing of LLM agents.
Golden-set replay and live-traffic scoring catch different things, and neither one substitutes for the other. Replay catches regressions against a known reference. Live scoring catches distribution shift: the traffic itself has changed shape rather than the pipeline breaking. A score drop on production traffic with no matching drop on the golden set almost always means the input distribution moved, not that the system got worse, and a team running only one of these two checks will misread which situation it's actually in.
The strongest version of this discipline treats every past regression as a future test case: production failures get folded into the evaluation dataset, the same scorers run inside CI, and pull requests get gated on quality thresholds before merge. That sequence turns every incident into permanent coverage rather than a one-time scramble. Running a scheduled replay of the golden set daily, or on every deploy, and tracking those scores over time, is the most reliable signal available for catching drift before it reaches users. Skipping that gate to move faster does not make the regression disappear. It just resurfaces later, as a fresh incident, and that incident has to travel the entire diagnosis cycle all over again.
Most teams that skip this step aren't being careless. They skip it because running it properly is genuinely expensive and slow, and without systematic evaluation in place, every prompt change becomes a gamble and every model upgrade demands a round of manual QA. This means teams avoid paying for validation precisely because of the cost that validation is meant to prevent. That's the irony sitting at the center of this stage: cutting the corner to save time is exactly how teams generate more total latency across the whole cycle, not less.
Latency taxes at each stage compounding into a structural ceiling on improvement velocity
None of these delays sit next to each other quietly adding up. Instead, they multiply. A failure caught at the validation stage doesn't just cost the time spent validating; it sends the team back to observation, and from there the entire sequence, observation, attribution, change, validation, runs again in full.
That compounding structure means a team's actual improvement speed is set by its slowest stage, not its fastest one. A team that builds excellent automated tooling for making changes but still attributes failures by hand hasn't meaningfully sped anything up, because the attribution bottleneck absorbs whatever time the automation saved elsewhere. Speed at one joint doesn't transfer to the joints around it.
This mirrors something already true inside the agent loop itself. Because model generation and tool execution run sequentially rather than in parallel, tool latency grows as a share of total time whenever decoding gets faster. The improvement cycle behaves the same way: its stages run in sequence, and total cycle time is governed by whichever stage is slowest, no matter how quickly the rest complete.
LLM evaluation is exactly at that chokepoint, right at the validation stage, the last gate before anything ships. Without automation there, teams are stuck choosing between slow manual QA and shipping without knowing what they've shipped, and neither choice actually breaks the ceiling. Notably, the ceiling isn't a model problem. Most agent incidents trace back to tool-call failures, context truncation, and runaway loops rather than to the model reasoning badly, so upgrading the model does nothing to shorten a cycle whose latency lives almost entirely in the harness and in the engineering process wrapped around it.
The practical leverage points for teams that want to move faster without shipping blind
Given all that, the instinct to speed up every stage equally is the wrong one. The stages aren't equal contributors, and treating them as if they were spreads effort thin across joints that were never the bottleneck to begin with.
Observation is the highest-leverage place to start, because every other stage waits on it. Instrumentation that captures tool calls, intermediate LLM calls, and span-level failures, not just final outputs, shrinks the gap between a failure happening and a team actually seeing it. That single change shortens the entire downstream sequence, because nothing else in the loop can start until observation delivers a usable signal.
Attribution is the second lever, and it responds specifically to structure rather than effort. Tools built around the model-versus-harness distinction, the ones drawing on frameworks like AgentRx, Who&When Pro, and the interaction-centric taxonomy work, turn a slow, judgment-heavy process into something closer to a repeatable check. That doesn't eliminate hard cases. It does cut down how often a team burns a full cycle chasing the wrong layer.
Validation is where the compounding effect is easiest to break, and also where teams cut corners most often. Golden-set replay run on every deploy, paired with continuous production scoring, catches regressions before they become the next incident someone has to diagnose from scratch. That pairing is what actually prevents a shipped change from turning into next week's mystery failure.
None of this collapses the loop into something instantaneous, and it shouldn't be expected to. The structural nature of the delay is the whole point. What changes is where the ceiling sits: a team that shortens observation, structures attribution, and refuses to skip validation isn't just moving faster at each individual stage. It's breaking the multiplicative effect that turns four separate delays into something worse than their sum.


