Production Learning Review

Human-in-the-Loop Review Gates for Agent Harness Changes

Treating harness changes as production infrastructure requires human sign-off like code review does.

Features Editor · · 12 min read
Cover illustration for “Human-in-the-Loop Review Gates for Agent Harness Changes”
Feedback Loop Architecture · September 30, 2026 · 12 min read · 2,671 words

Human-in-the-Loop Review Gates for Agent Harness Changes.

The harness as the deciding factor in agent reliability

Human-in-the-loop review gates for agent harness changes are a concrete engineering practice. They define when, where, and how a human has to sign off on a prompt, a tool schema, a workflow step, or a memory rule before it reaches production, and they apply the same rigor to that sign-off that any serious team already applies to a code change.

The shorthand "agent equals LLM plus memory plus tools plus planning plus action" has been repeated so often that it now hides more than it reveals. What a production harness actually governs looks a lot closer to runtime infrastructure than to prompt engineering: workflow design, loop engineering, evaluation, permission controls, and persistent state management.

The evidence for this is not abstract. Manus rewrote its harness five times in six months without changing the underlying model. A team building Deep Research re-architected the system four separate times in a single year, and the driver each time was workflow structure and context management. Vercel stripped 80 percent of the tools available to its agent and got better results, not worse. None of these are model stories. They are harness stories, and the January 2026 account of them makes the point: the engineering work that matters has moved.

The stakes of getting that engineering wrong are not small. An estimate cited in an analysis of AI harnesses puts the failure-to-production rate for agent projects at around 88 percent Best AI Harnesses to Supercharge LLM Models | Pinggy. Most of those projects are failing because the harness around the model cannot survive contact with real traffic, not because the model is weak Best AI Harnesses to Supercharge LLM Models | Pinggy. That is the premise this piece works from: if the harness is the surface where reliability actually gets decided, then changes to it carry the same risk profile as changes to any other production system, and deserve the same governance. Per the arXiv survey "Externalization in LLM Agents," capabilities are increasingly externalized into memory stores, reusable skills, interaction protocols, and the surrounding harness that coordinates them, with the model treated as a frozen reasoning calculator.

Harness changes and their risks

A harness has at least six distinct surfaces that teams modify, and each one fails in its own particular way. Prompts get edited, and a changed instruction can silently reshape tool-selection behavior across every downstream step, invisible to anyone checking only the final output. Tool schemas get adjusted: adding, renaming, or dropping a parameter changes what the model is capable of calling and how, which is a different category of failure from prompt ambiguity, not a milder version of it.

Workflows get restructured too — step order, branching logic, loop-termination conditions — and a change there can introduce an infinite loop or quietly skip an error-handling branch that used to catch a known failure mode. Memory rules change what persists across turns, so a constraint written into memory after a single human rejection in one session becomes binding for every future session, whether or not that was the intent. Evals themselves get changed, and a change to the quality gate can end up masking a regression rather than catching it, which is arguably the most dangerous failure on this list because it disables the safety net silently. And permissions — which tools are exposed, in what order, under what conditions — are an engineering choice with measurable consequences: Vercel's 80 percent tool reduction is the clearest illustration available that trimming the permission surface can improve reliability rather than constrain it.

These failures do not stay contained. A 2025 arXiv preprint on multi-agent failures found that 17.14 percent of failures are step repetitions and 13.98 percent are mismatches between reasoning and action Best AI Agent Evaluation Tools | AugmentCode. Both categories slip past any review that only checks the final output, because the failure is already embedded in the trajectory long before the run ends Best AI Agent Evaluation Tools | AugmentCode. A wrong tool call at step two corrupts every step that follows it, which is the structural reason a harness change cannot be treated like a config edit with local scope. It has blast radius by design.

Frontier labs control what might be called the inner harness, the foundational safety layers, native tool-calling, and raw context that ship with the model. The deploying organization engineers the outer harness, the feedback layer that handles output validation, tracing, observability, and human overrides. That distinction matters because it settles a question teams sometimes avoid: the review surface belongs entirely to the organization deploying the agent, not to the model provider. There is no vendor to defer to on this one.

What a review gate is

A review gate for a harness change is not a UX affordance, some polite pause screen designed to make an operator feel involved. It is an authorization checkpoint that has to clear before a changed prompt, tool schema, workflow step, or memory rule reaches production agents. The intent resembles pull-request review: a second set of human eyes examines a change before it ships. But the resemblance stops at intent, because a harness change asks the reviewer to evaluate non-deterministic behavior across a distribution of inputs, not to check a deterministic function against a spec.

A 2026 arXiv paper on human-in-the-loop agent oversight describes three structural intervention patterns. Pre-execution approval pauses the agent before a consequential action and requires explicit confirmation before it proceeds. Post-execution review lets the agent act first, then surfaces the result for inspection before it commits or moves downstream. Escalation triggers let the agent run autonomously under normal conditions but halt and request input the moment a specific risk signal fires — low confidence, sensitive data, an irreversible operation coming up. Hook systems generalize all three: operators attach validation checks or notifications to lifecycle events like a tool invocation, a file write, or a subagent spawn, which turns autonomy into a configurable harness parameter rather than a fixed property baked into the agent.

The deeper difference from ordinary code review is that the reviewer is not asking whether the logic is correct in isolation, but whether the change behaves correctly across the traces the agent actually meets in production. HAS-Bench, built by researchers at the University of Tokyo, University of Illinois Chicago, MBZUAI, McGill University, and Zhejiang University, formalizes exactly this point. Its benchmark shows human participation substantially improves task completion and failure recovery, but the gains depend specifically on when, how, and by whom the human input is exercised. A gate placed at the wrong point in the loop adds interaction cost without improving anything downstream.

Choosing the right gate threshold

Over-gating is measurable. HAS-Bench's A4 condition, maximum agent-assisted clarification, produced a 50 percent increase in turns compared to the A3 equal-partnership condition, with no corresponding gain in correctness Galileo. That is diminishing returns from excessive checkpoints, quantified rather than assumed Galileo. So the design question is never "should this be gated." It is "does this particular change clear a threshold that justifies the interruption."

Four criteria do most of the useful work. Irreversibility asks whether the change touches an action the agent takes that cannot be undone: a database write, an external API call, an email sent to a customer. Blast radius asks how many agent runs the change touches per day and how many downstream steps depend on the component being modified. Confidence signal asks whether prior traces already flag this surface as risky — a low tool-call success rate, recurring step repetition, latency creeping upward. And regulatory scope asks whether the change affects a decision class covered by rules like EU AI Act Article 14, which becomes enforceable on December 2, 2027 for Annex III high-risk systems, or by CFPB explainability requirements. These mandates make demonstrable human oversight a legal obligation for high-risk systems.

Replit's own product behavior is a clean illustration of reversibility-based calibration in practice: its agent generates code freely, no gate, but deployment is gated. Generation is cheap to throw away. Deployment is not, and the threshold reflects that asymmetry exactly.

Tool-authorization gates deserve their own category, separate from the four criteria above, because of where the enforcement sits. Policy evaluated inside the LLM's own context is visible to the model and can be routed around through prompt-level manipulation. Policy enforced at the tool execution layer, before the call is dispatched, means a blocked invocation simply never happens from the agent's point of view, and no reasoning context gets polluted trying to explain the block. That is a meaningfully stronger guarantee, and it argues for pushing authorization gates below the model rather than trusting the model to respect them.

The gate's own trigger logic is itself a harness configuration. It should be versioned alongside the changes it governs, not left as tribal knowledge in someone's head.

What the reviewer inspects

Handing a reviewer a diff of a prompt template or a revised tool schema leaves them with almost no basis for judgment. The change looks syntactically clean either way. What matters is how it behaves across the traces where the old version failed, and across the traces where the new one might newly fail, and neither of those is visible from a diff alone.

A review payload that is actually useful contains four things. The specific failure traces that motivated the change are the real trajectories. The layer attribution: which harness surface — prompt, tool, workflow, or memory — is identified as the root cause, and what evidence supports that attribution, since knowing a run failed tells a reviewer nothing without knowing which layer caused it. The proposed change is presented in a form that can actually be inspected. And replay results: how the proposed change performs against the historical traces that triggered it, and against a regression set drawn from traces that were already passing.

For tool-authorization gates specifically, the pending invocation payload needs to carry the full delegation chain, not just the identity of the immediate calling agent. The approver needs to see who delegated authority to whom, all the way back to the original triggering user. An approval that shows only the sub-agent's identity cannot satisfy SOC 2 or HIPAA audit trail requirements, and that gap separates a gate that produces a real audit trail from one that produces paperwork.

A survey out of TUM, drawing on 55 papers, found that existing observability platforms already capture and replay detailed agent traces, but leave developers to work out which step was responsible, why, and how to repair the run. That is precisely the gap a well-built review payload is meant to close.

One caution for teams leaning on LLMs to help with the review itself: when execution logs run long and distributed, an LLM acting as a review aid tends to settle on a plausible-looking failure before it has actually examined all the evidence Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures. The Continual Search framework addresses this by iteratively pushing the judge back toward unresolved evidence rather than letting it stop at the first coherent story Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures. On the TRAIL benchmark, that approach raised GPT-5.5's Weighted F1 from 0.386 under passive reconsideration Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures. Reviewers using LLM assistance should build prompts that demand full evidence coverage, or they inherit this bias without realizing it.

Replay validation as the mechanism that makes a gate meaningful

Synthetic benchmarks measure performance against curated inputs. Production traffic is a different distribution entirely, and a change that passes cleanly on a benchmark can still regress on the tail of what real users actually send an agent. This is the reason replay validation, not a benchmark score, has to sit at the center of a meaningful gate.

Replay in practice is straightforward to describe, if not trivial to build. Take the trace corpus, specifically the failure traces that motivated the change plus a representative sample of traces that were already passing. Run the proposed change against those same inputs and compare outcomes directly: did the failure mode actually disappear, and did anything that used to pass now break? Then gate on the result. If replay shows regression anywhere in that comparison, the change does not proceed, no matter how clean the diff looked going in.

Observability platforms that capture and replay detailed agent traces already exist and give teams the raw material this validation needs. But those platforms stop at capture and replay. Interpreting the result and deciding how to repair the run is still left to the developer. LangGraph's interrupt() primitive shows what the underlying infrastructure has to guarantee for any of this to work: the graph pauses at exactly the interrupt point, persists the full state snapshot to a checkpointer, and resumes only once a decision comes back. Without a checkpointer configured, there is nowhere for that state to live, and the pattern simply does not function. Replay-based gates carry the identical dependency: no durable state, no meaningful replay.

Once this infrastructure is in place, a reviewer's approval is no longer an act of faith in a diff but a conclusion drawn from observed behavior on real inputs. That shift, from trusting the diff to trusting the replay, is what separates a gate that does real work from one that exists for the audit log. A May 2026 paper on code as agent harness names regression-free harness improvement as an open challenge in the field. Replay against production traces is the practical answer available to teams right now.

Fitting review gates into the existing engineering workflow

None of this requires inventing a new discipline from scratch. Harness changes should be versioned and diffed exactly like code: prompt templates, tool schemas, workflow graphs, memory rules, and eval configs all belong in source control, under the same pull-request discipline as any application code.

Gate placement in CI follows a simple split. Automated replay runs on every harness change pull request, and a failure there blocks the merge outright. Human review only triggers when a change clears the criteria laid out earlier — reversibility, blast radius, confidence signal, regulatory scope — not on every single change that comes through. And when a change does reach a human, that reviewer approves or rejects based on the replay results and the root-cause attribution, not on the diff by itself.

The scale of the governance gap this closes is concrete. Enterprise environments now run something on the order of 82 autonomous agents for every human overseeing them, and only 22 percent of organizations treat their AI agents as identity-bearing entities with formal access controls Human-in-the-Loop Tool Calling: Approval Gates for AI Agents | ScaleKit. Gates close that gap systematically.

Gates should also learn from their own misses. When a production failure traces back to a harness change that shipped without adequate review, the post-mortem ought to feed directly back into the trigger criteria, tightening or loosening thresholds based on what actually happened rather than what a policy document assumed would happen. The "Who&When Pro" benchmark, covering more than 12,000 labeled trajectories across agent frameworks, domains, and modalities, underscores just how varied production failure patterns actually are Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures. No single gate design transfers cleanly across contexts, and teams should expect to tune their own trigger criteria against their own trace corpus rather than importing someone else's thresholds wholesale Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures.

The discipline compounds the same way code review culture compounds inside a maturing codebase. Every reviewed and replayed change adds to the trace corpus. A larger trace corpus makes the next replay set more informative. A more informative replay set makes the next gate a sharper instrument than the one before it. It is the same mechanism from software engineering, running on a different kind of system.

Sources

  1. 2025 Was Agents. 2026 Is Agent Harnesses. Here’s Why That Changes Everything. | by Aakash Gupta | Medium
  2. HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
  3. Code as Agent Harness
  4. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
  5. How to Build Human-in-the-Loop Oversight for AI Agents | Galileo
  6. Human-in-the-Loop Tool Calling: Approval Gates for AI Agents
  7. Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

More in Feedback Loop Architecture