Production Learning Review

LLM Evaluation Metrics for Production Agent Quality

Benchmark scores don't measure whether agents reach production safely.

Senior Staff Writer · · 11 min read
Cover illustration for “LLM Evaluation Metrics for Production Agent Quality”
Online & Offline Eval · October 9, 2026 · 11 min read · 2,565 words

An agent that clears a benchmark suite can still break the first week it handles real traffic, and the reason is structural: benchmark scoring measures whether an output looks right, not whether the agent got there sensibly. The evaluation methods most teams carry into production were built for models that produce a single output from a single input. Agents don't work that way, and the mismatch between how they're tested and how they fail is the subject of this piece.

Benchmark scores after an agent ships to production

Single-turn scoring and held-out benchmark accuracy only check one thing: whether a model's output matches what a grader expects. They say nothing about how the agent arrived at that output, and for an agent, the path matters as much as the destination. An agent can make three redundant API calls, hallucinate a tool it can't access, and retry a failed step five times, and it can still land on the correct final answer. Output-only scoring gives that run a passing grade. It never measures what the run cost, how many steps it wasted, or how close it came to failing.

That blind spot has a direct financial cost. Two agent implementations can post nearly identical accuracy numbers on the same benchmark task while one costs several times more to run, because the waste lives in the trajectory, not the output, and trajectory-level waste is invisible to a scorer that only looks at the last message. A team comparing two agent designs on accuracy alone has no way to see this difference. They'd ship the more expensive one without knowing it, because the metric they used wasn't built to catch it.

Benchmark environments compound the problem because they test under conditions production never offers. Benchmarks run on clean inputs, predictable tool responses, and controlled sequencing. In production you deal with ambiguous requests, third-party APIs that time out or rate-limit without warning, data that arrives in undocumented formats, and inputs designed to probe for weaknesses. The gap between how an agent performs in a lab setting and how it performs once real traffic hits it is well documented, and it's large enough that a benchmark score tells you almost nothing about how the system holds up once it's live.

Evaluating an agent versus evaluating a model

Agents are not models with extra steps. An agent is a frozen model wrapped in a harness: a set of prompts, tools, memory stores, workflows, and skills that shape every decision the model makes before it generates a single token. Evaluating an agent means evaluating that entire wrapper, not just whatever the model produces at the end of it.

That distinction matters because most production failures start in the harness, not the model. When a prompt encodes a bad rule, when a tool is called with a schema that no longer matches what the tool expects, or when memory returns stale context from an earlier session, the model's output can look broken even though the model behaved exactly as instructed. Teams that trace a failure back to "the model" without checking the harness are routinely wrong about where the fix belongs, and a fix aimed at the model when the problem lives in the prompt or the tool schema won't hold.

Agent tasks also run in multiple steps, and each step is a place where error can enter and propagate. A wrong tool call or a hallucinated argument passed into a function can corrupt everything downstream, even when the final generation reads as fluent and confident. A model evaluated on a single input and a single output has no equivalent failure mode: there's nowhere for an early mistake to hide, because there is no "early." An agent's entire value and entire risk live in the sequence of decisions between the first input and the last output, and that sequence has to be the unit of evaluation, not the endpoint.

The five dimensions production agent evaluation must score simultaneously

No single metric captures whether a production agent is working. Five dimensions need to be scored together, because optimizing for any one of them alone creates a blind spot that eventually causes an incident.

Task completion rate comes first, and it's less binary than it sounds. An agent can finish part of a job, or finish the wrong job while it sounds sure of itself, or finish the right job but leave side effects nobody asked for. The worst case in this group is an agent that fails but believes it succeeded: it throws no error, sends no signal that anything went wrong, and quietly corrupts whatever system receives its output. Catching that case requires defining, in advance, what counts as done, what counts as partial, and what counts as failed while convinced otherwise.

Accuracy needs to be checked at more than one level. Surface accuracy, whether an output is syntactically well-formed, and semantic accuracy, whether it's logically correct, are separate questions. A code-generation agent can produce code that runs with no errors even when it does the wrong thing. Accuracy for an agent should be broken into separate pass/fail checks: did it extract the right information, did it pick the right tool, did it format the output correctly, did it handle the edge case without guessing.

An agent's hallucination rate needs its own line, apart from general accuracy. An agent hallucination carries more risk than a chatbot hallucination, because the agent acts on what it invents: it calls an API endpoint that doesn't exist, references a database table that was never created, or files a report built on numbers nobody generated. Hallucination has to be measured at the trajectory level, step by step, because a hallucinated tool call three steps into a run can corrupt everything that follows even if the final message the agent returns reads as entirely plausible.

Cost per task is the dimension most teams underprice. Total cost includes every LLM call across every step, plus whatever the agent spends calling tools and external APIs, so you need to track it step by step if you want to find which parts of a run are burning money disproportionately. You need to track cost at the trajectory level to catch the case where two agents land on the same accuracy number while one costs far more to run. When cost climbs across releases without a matching gain in accuracy, that's a sign of trajectory degradation, more retries, more redundant calls, not a sign that the underlying tasks got harder.

Reliability closes the list, and it splits into three parts under the ReliabilityBench framework: consistency, whether the agent returns the same result across repeated runs of the same input; robustness, how much performance degrades when a task is rephrased or an input gets noisy; and fault tolerance, whether the agent recovers when an API times out or a rate limit kicks in. Consistency failures are the hardest to catch because a single test run can't reveal them. An agent that succeeds once and fails twice on the identical input looks fine under a one-shot evaluation and only reveals the problem once it's run several times over.

The three structural failure categories that production traffic exposes and benchmarks cannot

Three failure patterns account for most of what breaks in production, and all three share a property that benchmark evaluation can't reach: they depend on state, on sequence, and on environment, none of which a static output check can see.

Tool schema drift happens when a tool's expected input changes, a field gets renamed, a new property becomes required, a flat key turns into a nested object, an enum changes its allowed values, a validator gets stricter, while the agent keeps calling it under the old contract. The correct response to a schema mismatch is not to retry with the same arguments. The tool itself is broken in that case, and an identical retry produces an identical failure every time. Agents without explicit schema validation logic have no way to tell a transient network hiccup apart from a structurally broken response, and they'll keep retrying a call that was never going to succeed.

Prompt ambiguity is a harness problem, not a model failure, separate from hallucination. In multi-turn tool-calling, every intermediate tool output gets added to the conversation history, and if the agent pulls tools dynamically based on whatever the latest scratchpad state looks like, each turn's accumulated context can bias the next tool choice a little further from the original goal. That's a semantic drift loop, and it originates in how the prompt and retrieval logic were designed.

Workflow loops are the third category, and they have a detection heuristic simple enough to run in production: when an agent calls the same tool with the same arguments three or more times in a row, that's a strong signal it's stuck. Loops are expensive in a specific way, because every cycle adds token cost and latency without moving the task forward, so cost-per-task tracking is one of the fastest ways to catch a loop before a user notices the agent never finished.

All three categories, schema drift, semantic drift, and workflow loops, share the same diagnostic requirement. None of them appear in a single output snapshot. A loop, a schema mismatch, or a prompt with no exit condition only becomes visible across a sequence of steps, and that's exactly the kind of signal a benchmark, built around single inputs and single outputs, has no mechanism to capture.

Trajectory-level analysis for root-cause attribution

Diagram: Instruction Compliance: Compaction Drop vs. Model Failure. Visualizes: Show the stark contrast between two failure modes that look identical from the outside but have opposite causes.

Knowing that a run failed is not the same as knowing why. An outcome-level signal tells you a run broke somewhere, but it carries no information about which layer of the system caused it, and treating that as a question you can answer by inspecting the final output is where most production debugging stalls out.

The fault in an agent run has to be assigned to a specific responsible component: the model, the prompt, a tool, the workflow logic, memory, or the surrounding environment. The intervention changes completely depending on which one it is. A harness failure needs a harness fix. A model failure needs something else entirely, and applying one when the other is needed leaves the actual cause untouched.

A documented case makes the stakes concrete. In a long-running agent session, an agent can ignore an instruction a user gave earlier in the conversation for two entirely different reasons: the harness's context compaction may have dropped that instruction from what the model sees, or the instruction may still have been present and the model simply failed to follow it. One published study examining this exact scenario found 31.8% compliance when the constraint had been dropped by compaction, against 99.7% compliance when the constraint remained available to the model. The two failures look identical from the outside: the agent didn't follow the instruction. Only a trajectory-level trace, one that shows what the model actually saw at that step, reveals which of the two actually happened.

Root-cause attribution, under this framing, is a search problem across a long trajectory, not a single judgment made after the fact. The relevant evidence is often sparse and spread across actions that happened many steps apart, so a one-shot classification pass over the final trace misses it, the way a single glance at an x-ray misses a hairline fracture. A benchmark built for this problem, LongRCA Bench, scores two separate targets rather than one: responsible-role attribution, which component caused the failure, and root-cause step localization, which specific step in the trajectory was the earliest decisive cause. Treating those as one combined score hides the cases where a team correctly identifies that the harness was at fault but points to the wrong step, or correctly finds the step but misattributes the layer responsible for it.

Without a trace that links every action to the prompt state that produced it, the tool call that followed, and the value that call returned, attribution has no evidence to work from. What's left is guesswork, and guesswork tends to fix the symptom while the actual fault lives undisturbed in the layer beneath it.

What production traces capture beyond offline evaluation

Conventional logs can't reproduce a non-deterministic LLM failure after the fact. Span-based trace replay has become a baseline requirement for any team that wants to diagnose a production failure.

A complete production trace holds far more than a log line. It captures the user's original input and the agent's final response, every intermediate LLM call along with the exact prompt and completion for each one, every tool call with its name, its arguments, and what it returned, token counts and latency measured per step, the agent's planning and reasoning as it moved through the task, and any retrieval results pulled in along the way. That data is what turns each of the five evaluation dimensions from a theoretical target into something a team can actually measure. You need per-step token and latency data for cost per task, so you can find the expensive step. Hallucination rate needs the intermediate tool calls, not just the final message. You need multiple full traces of the same input to check reliability's consistency and robustness against each other. None of the five dimensions can be scored from a final output alone, and the trace is what supplies the steps in between.

Traces also enable a kind of evaluation a single final-answer score can't perform: scoring each turn on its own, as its own classification problem. Each turn can be labeled for a specific event, whether the agent is looping, whether it's violating a stated policy, whether the user's tone signals frustration, and these per-turn labels feed directly into the reliability dimension described earlier.

The way these traces get scored has moved as well. Reference-based metrics compare an output against a known correct answer, but they run into limits fast on multi-step agent trajectories, because there often isn't one correct path to a given outcome. The shift has been toward using an LLM as judge for trajectory-level quality, scoring the sequence of decisions on its own merits. And when a production regression gets converted into a test case, it adds a permanent entry to the offline evaluation suite, turning it from a static artifact of initial development into a running map of every failure mode the agent has actually encountered.

Replaying traces before shipping changes to validate improvements without shipping blind

If you ship a fix without checking it against real production traces, you ship blind. A fix can resolve the synthetic test case it was built against while leaving the actual production failure mode untouched, or introducing a new one nobody tested for, and the only way to know which happened is to run the fix against what really went wrong.

That's the function an evaluation harness serves: the layer between running a few evals once during development and having live traces flowing in from production. It gives a team a repeatable way to define golden tasks drawn from real incidents, so you can replay tool calls deterministically and test a fix against the exact sequence that failed before, score each run against a rubric built for the task, and block a regression in CI before it reaches a user. Built around production traces rather than synthetic inputs, that loop turns agent evaluation into a standing discipline that catches the next trajectory failure before it ships.

Sources

  1. Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
  2. LongRCA Bench: Root-Cause Localization in Long-Horizon Agent Trajectories

More in Online & Offline Eval