When Offline Evals Are Insufficient for Production Agent Changes
Offline evals can pass while production agents fail in ways frozen test sets never see.

Offline evaluation exists to answer one question cheaply and reliably: did this change make the agent worse on the cases the team already knows about? That is a narrow job, and it is worth doing well before it is asked to do anything more. The unit of work is a versioned trajectory dataset, an ordered record of system prompt, user input, tool calls with their arguments and returns, intermediate reasoning, and final response, each scored by a rubric written in code against a known expected outcome. Good offline suites break that score into per-dimension assertions, checking tool selection, argument extraction, result use, error recovery, plan coherence, and task completion separately, so a weak axis cannot hide behind a healthy average. This setup's value comes from its reproducibility: fixed inputs, known answers, a test that runs the same way every time it runs, which is what a pre-release gate needs. It belongs in CI on every prompt change, every tool schema change, every model swap. None of that is in question. What offline evaluation gives a team is a floor, a minimum bar below which a change should never be allowed to ship, and everything that follows in this piece is about what sits above that floor, where the floor cannot reach.
Why "passing evals" and "ready for production" differ
An offline eval is a frozen bet on three things: what users will ask, how tools will behave, and which prompts will still be running by the time anyone looks at the results again. Every one of those three starts drifting the moment the eval set is written down. Offline evaluation can confirm the agent did not regress on yesterday's tasks. It cannot anticipate an input nobody wrote a test for, a multi-turn conversation that wanders somewhere the test set never modeled, or a tool failure that only happens against a live system under live load. The comparison that holds up here is a unit test left in place after the function underneath it gets rewritten: the test still passes because it never knew enough to object, not because the code is still correct. An offline eval goes stale the same way, through no carelessness on the part of whoever wrote the rubric.
Most teams still score task completion by reading the final message and deciding whether it sounds right. A support agent can tell a user their refund was processed, their flight was rebooked, their ticket was closed, and the text can read as confident and correct while nothing downstream actually happened. Two traces can carry the same status code, the same latency, the same token count, and one of them resolved the user's problem while the other quietly failed it. A rubric that checks only the last message cannot tell these apart, because it was never built to look at the path, only the destination. This is why benchmarks like tau-bench verify database state alongside output content rather than trusting the final text alone: an agent that claims it booked a flight has to have actually booked the flight, and checking the claim instead of the ledger behind it is how "passed the eval" and "ready for production" come apart.
The six drift modes that structurally age every eval set
Eval sets do not fail all at once. They age through six distinct mechanisms, each working on its own clock and each exposing a different layer of the agent harness that a fixed test set was never built to reach.
Dataset drift is the slowest and the easiest to miss. The eval set reflects the team's best model of user intent at launch. Real users, given time, find intents nobody wrote a test case for. Offline scores stay flat because they are still measuring the old distribution, while production complaints diversify in a direction the test set has no visibility into, and most of what users report cannot even be reproduced against the frozen dataset.
Tool-API drift is less intuitive and does more damage before anyone notices. CI mocks the tool call, and the mock returns whatever payload it was built to return, on schedule, forever. The real endpoint, meanwhile, changes its schema, changes its error format, adds a rate-limit header nobody asked for, and the agent responds by retrying, timing out, and producing a plausible-sounding answer that was never grounded in a real tool result. CI stays green through all of it, because the mock has no way to know the vendor changed anything. The failure is quiet in a specific way: tool-call latency and retry counts climb in production, cost per successful task creeps upward, and the per-response rubric keeps passing the whole time. The judge grades against a contract that stopped being true, not by accident.
Retrieval-corpus drift works on documents. The index gets frozen at the moment the eval set is built. Months later the corpus has doubled, the chunker has been re-embedded, and the same query now surfaces a different set of top-k chunks than it used to. The generator does its job faithfully, grounding its answer in whatever it was handed, so groundedness scores stay healthy even as the answers themselves get worse, because faithful grounding in the wrong chunk is still faithful grounding.
User-distribution drift is a problem of input shape, not tooling. Hand-written test inputs look like what the team expected users to type. Real traffic arrives with different phrasing, longer context, tool-call sequences the test authors never imagined, and no amount of tuning the harness fixes a mismatch that lives in the data itself.
Agent-step compounding is structural, present from day one. A high per-step success rate looks comfortable in isolation. Multiplying it across a dozen steps in a long trajectory drops the end-to-end completion rate to something far less comfortable, a fact single-turn rubrics were never built to catch because they have no concept of cumulative error. Step and loop counts are invisible in a trace total too: an agent that looped eight times and one that answered cleanly on the first try can both return a single final answer, and the difference appears only when cost is attributed per step rather than per turn.
These six modes don't share a timescale, and that is part of what makes them hard to catch with one fix. Dataset drift and user-distribution drift accumulate over weeks of real usage. Tool-API drift and prompt drift can land overnight, the moment a vendor ships an update. Retrieval-corpus drift stays invisible until the next re-index. Agent-step compounding never arrives at all, because it was built into the architecture from the start. None of the six are solved by writing more test cases. A bigger frozen snapshot is still a snapshot.
Why multi-step agent traces make attribution harder than a failed rubric score suggests
In a single-turn system, a wrong answer and the cause of the wrong answer live in the same place. In a multi-step agent trace, they rarely do. A weak plan written at step one, a wrong tool call made at step three, a bad assumption formed early, all of these cascade through every step that follows, so the failure a rubric sees at the end is frequently several steps removed from the actual mistake. A constraint violation can surface late in a trace while its root cause was an orchestrator omission or an executor's dropped instruction much earlier, and without tracing causality across the full sequence of steps, there is no way to separate the root cause from the symptom it produced downstream.
Research on multi-agent root-cause analysis backs this up directly: the most common pitfalls, hallucinated data interpretation and incomplete exploration, persist across models regardless of capability tier, which points to these failures coming from the shared agent framework itself. That finding matters because it reframes what a failed trace is actually telling a team. If the error pattern occurs at the same rate whether the model is strong or weak, the fix was never going to be a better model.
This is where a lot of post-incident response goes wrong. A team sees a production failure, traces it to a bad final answer, and upgrades the model, assuming a smarter model will stop making the mistake. If the real fault lives in the prompt, in a tool's schema, or in the workflow logic that strings steps together, the model upgrade leaves the actual cause untouched and the failure mode returns under a different disguise. An offline eval score, whether aggregate or broken into per-dimension numbers, tells a team that a run failed. It does not tell them which layer of the harness, prompt, tool, memory, or workflow, is where the fix belongs. Knowing a run failed and knowing what to change are two different kinds of information, and offline scoring only ever delivers the first one.
What production traces expose that offline rubrics structurally cannot
Two traces from a support agent can look identical to an observability platform: same status code, same latency band, same token count, no errors logged anywhere. One of those two traces resolved the user's actual problem. The other told the user something was handled when it wasn't, and the user found out later, somewhere the observability platform was never watching: an observability layer records what happened mechanically. It has no way to record whether the agent did what the user actually needed.
Closing that gap means scoring live traffic, not only collecting it. Online evaluation applies the same kind of rubric used offline, but to real requests as they happen, catching the long tail no dataset was ever built to contain: drift, novel failure patterns, users who are visibly frustrated, tools misbehaving in ways nobody anticipated. The same rubric, attached as a score on live OpenTelemetry spans, server-side and after export, produces a regression signal the offline set structurally cannot produce, because the offline set is a snapshot and the trace stream is not. The eval surface moves from a fixed file sitting in a repository to something sampled and scored continuously as traffic actually arrives.
Tracing by itself is necessary but not sufficient here. It is the backbone that shows where a metric broke and surfaces failure modes the team has no name for yet, but a trace with no score attached shows only that something happened, not whether it happened well. Per-turn evaluation on live traffic catches categories of failure a final-answer check will always miss: a policy violation, a jailbreak attempt, a leaked system prompt, a moment where the user's tone turns from neutral to frustrated. None of these change the status code. None of them touch latency or token count. All of them are invisible to a log line and visible only to a rubric reading the actual content of the turn.
The traces that matter most are the ones that feed back into the offline set. Clustering the hardest failing production traces and promoting a representative sample into the dataset on a regular cadence, say weekly, is what keeps the gap between "passed offline" and "failed in production" from calcifying into a permanent blind spot. Without that promotion step, the offline set keeps testing the same fixed slice of reality while production keeps moving somewhere else.
What must supplement offline evals before shipping
Closing the structural gaps described above takes three things working together: the same rubric running in both the offline gate and on production traces, a feedback loop that promotes real failures into the offline dataset, and replay-based validation against actual recorded traces before any harness change gets committed.
The loop runs in a specific order, and skipping any stage reopens the gap. Offline eval feeds the CI gate. The CI gate protects what ships. Production trace eval scores live traffic with the same rubric. Trace eval clusters the failures it finds and promotes representative cases into the offline dataset. Optimization work runs against that expanded dataset, still under the same rubric. Winners from that process ship back through CI, closing the circle. Pulling any one stage out of that sequence breaks the loop, and the agent drifts again, quietly, the same way it did before the loop existed.
Replay validation is what makes this loop trustworthy. Recording real tool outputs and replaying them through a sandboxed version of the harness lets a team confirm that a proposed fix actually improves behavior on the traces that motivated the fix in the first place, not just on a synthetic case someone wrote afterward to feel good about the change. Shipping a harness change without replaying it against the traces that exposed the original problem means shipping a fix with no evidence it fixes anything.
Drift itself needs to be checked for on a schedule, not assumed away. That means sampling recent production traces at a regular cadence, re-running the agent on those same inputs inside a sandbox, and comparing today's outputs to what the agent produced on those same inputs before. Tracking tool-call distribution, schema conformance, cost, latency, and output content over time turns drift from a surprise into a measured trend.
None of this works without attribution specific enough to act on. Knowing a run failed is not the same as knowing whether the fault sits in the prompt, the tool schema, the workflow logic, memory, or the model itself, and a fix aimed at the wrong layer leaves the real cause sitting exactly where it was. The harness is where almost all of this improvement actually lives: prompts encode the standing behavioral rules the agent follows, tools expose external services along with their action schemas, invocation formats, and validation rules, memory carries forward prior observations and task outcomes, and skills package up reusable procedures. All four of these can be changed, tested, and shipped without touching the underlying model, and all four are only addressable by a team that can actually read the trace where things went wrong.
Offline evaluation remains the floor. It is the gate that proves a change did not make things worse on the cases already known, and it belongs in CI without exception. The architecture described here, live trace scoring, the feedback loop, replay validation, scheduled drift checks, specific layer attribution, is what has to sit above that floor before any agent change is trusted with real users.


