Replay-Based Eval for Validating Agent Changes Before Shipping
Real production traces catch multi-turn regressions that static benchmarks cannot see.

A pre-ship test suite can pass every case in the harness, but the agent it ships can still fall apart in production within a week. That gap exists because conventional offline evaluation checks an agent against tasks the team already knew about, and production never confines itself to what a team already knew about. The fix is a different validation surface: real production trajectories, replayed and scored, standing in for the synthetic scenarios that pre-ship testing has relied on.
Why pre-ship validation fails without production trajectories
A held-out dataset written last month is a frozen snapshot. It can tell you whether the agent still handles the inputs the team anticipated when the dataset was built, but it cannot contain the inputs, the tool schema states, or the multi-turn conversation drift that production will generate next week. The dataset is static by construction. Production is not.
Single-turn evaluation compounds the problem. The gap between per-step and end-to-end performance widens with every additional step in the chain, and a test harness built around single-turn cases has no way to see it, because it never runs the full sequence.
Three structural gaps separate offline benchmarks from production, and none of them close no matter how large the benchmark grows. Production requirements arrive as loosely worded business descriptions, not clean task specifications, so the agent often has to solve a problem that was never fully written down. Domain experts judge quality, but their standards shift over time, so what counted as a correct answer six months ago may not count today. And the evaluation itself has to keep pace with live traffic: a benchmark fixed at launch time is already behind by the time it ships. A team can write a thousand more test cases and still miss all three.
What a production trace contains as a validation surface
A production trace is a record of everything that happened during one agent run: the system prompt, the user input, the agent's reasoning, every tool call along with its arguments and return values, retrieval results, any intermediate LLM calls, the final response, and outcome metadata tied to how that response landed. The unit of analysis is the full trajectory, not a simple pairing of what went in and what came out.
That distinction matters because a trace preserves the exact conditions that produced a failure: the tool schema as it existed at the moment of the call, the state of memory at that point in the conversation, the routing decision that sent the request down one path instead of another. A written test case can approximate those conditions. A trace recorded them as they actually occurred, with nothing reconstructed after the fact.
The trace also captures what can be called harness state: everything that shapes how an agent behaves without touching the model's weights. All of these can shift between one release and the next, and a trace is the only record that shows what they looked like at the moment a given run happened.
Production inputs also tend to be messier than anything a benchmark author would construct by hand. A curated benchmark dataset cannot reproduce that mess. A real trace records it automatically because it is a record of something that actually happened, not something written to resemble it.
How failing traces become a regression harness
Replay eval takes a failing trace from production and turns it into a permanent fixture of the test suite. The mechanism runs in three steps. A failing trace gets promoted into the trajectory dataset. The agent change under review gets replayed against that same trajectory, reproducing the same inputs and conditions the original run faced. The output gets scored against the same rubric used both in continuous integration and in live traffic, so one scoring vocabulary carries across all three stages without translation loss between them.
Incident.io's AI SRE product, called Investigations, offers a concrete working example of this pattern. It replays historical incidents to its own agent and checks whether the agent finds the correct cause, a kind of time travel, since you already know the real resolution. That setup catches two things a live test cannot: whether the agent's understanding of an incident lags behind or runs ahead of what a human responder would conclude, and whether the agent would hallucinate a fix that was never actually possible given what was known at the moment the incident occurred.
Replaying the highest-volume user conversations at every release catches a category of regression that single-turn tests cannot touch: failures that only emerge across a sequence of turns, where no individual step looks wrong but the conversation as a whole goes off course. Shadow mode is a complement to replay, not a substitute for it: replay tests against history that already happened, shadow mode tests against the present without the risk of acting on it.
What turns this from a one-time check into an actual regression harness is cadence. Failing traces are promoted into the dataset on a regular schedule, so each new release has to face a dataset that has already absorbed every failure pattern production has surfaced up to that point. The full loop closes like this: offline eval feeds the CI gate, the CI gate feeds production trace evaluation, trace evaluation feeds an error feed that clusters failures and promotes representative cases into the dataset, optimization runs against that expanded dataset, and the winning changes ship back through CI. Removing any one stage from that loop stops it from closing. An agent left on an open loop drifts within weeks, because nothing is feeding new failures back into what gets tested.
Why the score must be per-layer to be actionable
An aggregate score on a replayed trajectory hides more than it reveals. A healthy-looking overall number can sit on top of one severely failing layer, and a failing overall number gives no indication of where to look next.
Consider a trajectory where tool selection scores well but argument extraction scores badly. Averaged together, the result can still look acceptable, even though the production failure is riding entirely on the argument-extraction layer. A team chasing that failure with only the aggregate score to go on can spend days bisecting the system's tool-selection logic before finding that the actual fault lies in argument extraction.
Per-dimension assertions in CI solve this directly. Instead of a single pass or fail, the gate checks each scoring dimension on its own and exits with a non-zero status on whichever dimension fails. That failure points straight at the harness component responsible.
AlphaEval's design makes the same case at a broader scale. Its evaluation spans multiple distinct paradigms within each domain it covers, including LLM-as-judge scoring, reference-driven metrics, formal verification, rubric-based assessment, and automated UI testing. Its central finding is that scaffold, the surrounding harness and tool infrastructure around a model, matters as much as the model itself. An aggregate, model-level score cannot see a regression that originates entirely in the scaffold, because you never isolate what the scaffold contributes from everything else being measured.
None of this works if the scoring vocabulary shifts across stages. The rubric used in the offline dataset, the one applied to live clusters, the one attached to each new dataset entry, and the one driving optimization all have to be the same rubric. If the definitions shift between any two of those stages, the loop built in the previous section stops functioning, because a score computed one way can no longer be compared against a score computed another way.
The six failure categories replay scoring must distinguish
Agent failures sort into five broad categories: planning errors, tool errors, retrieval errors, reasoning errors, and safety or policy violations. The harness, the surrounding infrastructure rather than the model itself, accounts for the majority of these. Replay scoring that only looks at the final output misses the actual root cause in most cases, because the failure originated somewhere upstream of what the output shows.
Tool schema drift is the most dangerous of these in practice, precisely because it produces no obvious crash. None of these generate a clear error signal on their own, so they tend to be the last thing diagnosed when scoring happens only at the aggregate level.
Planning failures take three recognizable shapes, and you can see all three directly in a well-built replay score. An agent with no format specification can produce a correct answer that is still unusable, because the downstream step cannot parse it.
Multi-agent systems shift the distribution of these failures. The MAST failure taxonomy, built from an analysis of more than 1,600 execution traces presented at NeurIPS 2025, found that specification problems, meaning role ambiguity, unclear task definitions, and missing constraints, make up the largest single category of failure. Inter-agent misalignment follows close behind. So that finding points teams running multi-agent systems toward the specification layer first, instead of assuming failures originate in any individual agent's reasoning.
If a replay harness scores only final task completion, it cannot tell a planning loop apart from a tool argument error, or either of those from a schema drift problem. Scoring at the level of the individual layer turns a vague finding, "this trace failed," into something a team can act on directly: argument extraction regressed on date strings along the booking path.
Root cause attribution: turning a scored failure into a directed fix
The step where a failure becomes visible is almost never the step that actually caused it. A score that names the failing layer gets a team partway there, but it still leaves the harder question open: which specific change, at which specific point in the trajectory, produced the bad outcome.
An outcome-level signal tells a team only that an execution failed. Root cause attribution turns that into something usable by tracing back to the first point where execution diverged from what a correct run would have done, and assigning responsibility to the component actually at fault, whether that is the model, the harness, the environment, or the grader itself.
That assignment determines what happens next. AlphaEval's work identifies six production-specific failure modes that only become visible once evaluation covers the complete agent product rather than isolated model calls in isolation, reinforcing that attribution has to reach the scaffold and not stop at the model boundary.
Long execution traces introduce a separate risk during attribution itself. As the logs being reviewed grow larger and spread across more components, a judge model reviewing them can settle on a plausible-looking root cause before it has worked through all the evidence available. If diagnosis proceeds iteratively, revisiting the evidence across several passes rather than accepting the first plausible explanation, it surfaces root causes that a single review would miss.
The output of this process should be a ranked list of candidate fixes, each attached to the evidence that supports it, rather than a fix applied automatically without explanation. That keeps the decision in front of the engineer, who can inspect the reasoning behind a proposed change before deciding whether to ship it.
Shipping a harness change without replay validation
A diagnosis is an input to the next step, and the next step is validation: checking a proposed harness change against the traces that already proved a given failure can happen. Skipping that check, and validating only against the single new trace that surfaced the problem, all but guarantees the fix will regress on a pattern the team has already encountered once before.
If a team never replays the fix against the full set of historical failures, skipping that replay lets the regression stay invisible until it reaches production again, and then the team has to debug a problem it already solved once.
Replay validation against historical traces is the only pre-ship check that measures a proposed change against failures you already proved can occur in this specific system. The traces recorded the actual conditions directly.
The same rubric has to score the validation pass that scored the original failure. The alternative to replay validation, manual log review guided by developer intuition, does not scale as trace volume grows, and it cannot systematically show whether a fix introduced a new failure class while it closed the old one.
The full loop of replay validation as an engineering discipline
When replay-based validation sits inside the full improvement loop, offline evaluation, CI gate, production trace evaluation, error feed, optimization, and re-validation, each release gets tested against a regression harness that keeps growing more representative with every incident production surfaces. A team working inside that loop can name the failing layer, the change that fixed it, and the evidence behind that fix, for every release they ship.
The first stage is offline evaluation: a versioned trajectory dataset, stratified by tool, by argument edge-case bucket, and by error code, scored across all six dimensions using code-defined rubrics. Failing production traces get promoted into that dataset on a fixed cadence, so the dataset tracks what production is currently doing.
The second stage is the CI gate: per-dimension threshold assertions run on every pull request, and the gate exits non-zero on whichever axis fails. Because the CI gate uses the same rubric as the production scorer, a regression caught here predicts a block that would otherwise occur at runtime. The loop that follows, production trace evaluation feeding the error feed, the error feed promoting representative failures into the dataset, optimization running against that expanded set, and the winning changes shipping back through CI, is what keeps a team from relearning the same failure twice.


