Production Learning Review

Building an LLM Eval Pipeline in GitHub Actions

Tiered checks and statistical gates replace expensive single-judge sweeps on every PR.

Contributing Editor · · 11 min read
Cover illustration for “Building an LLM Eval Pipeline in GitHub Actions”
Online & Offline Eval · October 10, 2026 · 11 min read · 2,489 words

A tiered cascade, not a single LLM-judge sweep, is what makes an eval pipeline survive contact with a real team's pull request volume. At 4:47 PM on a Friday, a gate built on one frontier judge fires green after running every rubric against every example, the PR merges, and within hours the deployed agent regresses sharply on a traffic class the dataset never covered. The suite held because the question was never asked, not because the agent was healthy.

Why a single LLM-judge check on every PR fails

The setup that most teams reach for first looks reasonable on paper: one LLM-judge sweep, a small curated dataset, run on every pull request. In practice it produces a gate that fails in one of three ways. It costs too much to survive a real team's PR volume, it runs too slowly for engineers to wait on before merging around it, or it operates on a sample too small to separate a genuine regression from ordinary judge noise. Few teams get all three right by accident, because the three properties pull against each other.

Cost per PR, time to verdict, and statistical significance form a triangle where improving any two degrades the third. A gate with a tiny dataset can run fast and cheap, but its confidence interval is wide enough that a real quality drop is comfortably inside the noise band, and the gate starts firing on nothing, training engineers to ignore it. A gate can pair a large dataset with a frontier judge on every example and get a defensible verdict, but the eval bill outpaces the model API spend within months, so someone quietly disables the check instead. A gate that tries to be fast and cheap and still statistically sound by trimming the judge's reasoning depth just produces a verdict nobody trusts. Sophistication in the judge itself does not rescue a design that ignores one leg of this triangle. A frozen-baseline, floor-only threshold has the same blind spot: it catches a catastrophic drop but misses the slow drift that accumulates over weeks of small prompt and tool changes, because nothing ever crosses the floor on any single day.

Diagram: The Cost-Speed-Significance Triangle. Visualizes: Illustrate the fundamental tension in LLM eval gate design: three properties — cost per PR, time to verdict, and statistical significance — form a triangle where improving any two degrades…

How a tiered cascade resolves the cost-speed-significance triangle

Diagram: How the Three-Tier Cascade Routes Each Example. Visualizes: Show the decision flow of the tiered cascade: every PR example enters Tier 1 (deterministic checks: JSON schema validation, regex, NLI-backed classifiers on CPU — resolves the…

A tiered cascade resolves the triangle by changing what runs where, not by making any single check faster or cheaper in isolation. The cheapest, most deterministic validators run first, on every PR, and only the subset of cases they cannot resolve gets escalated to a classifier, and only the subset the classifier flags as uncertain gets escalated again to a frontier judge. The expensive step still happens, but it applies to tens of examples instead of hundreds, and the dataset's full statistical weight is preserved because nothing upstream silently drops a case, it only defers it.

Think of the three tiers as a decision tree rather than three parallel options competing for budget. The deterministic tier handles schema validation, regex checks on tool-call structure, and NLI-backed classifiers run on CPU, resolving the clear majority of cases without any call to a model API. Tier 2 is a frontier LLM judge, but it only sees the examples Tier 1 flagged as low-confidence, typically a fraction of the full dataset, scoring rubrics like groundedness, context adherence, completeness, task completion, and function-calling accuracy. Tier 3, covered later in this piece, is the statistical gate that decides if Tier 2's verdict is a real regression or just ordinary variance.

One routing decision deserves explicit attention because getting it wrong defeats the entire cascade: rubrics that require judgment from the start, such as helpfulness, tone, or brand voice, should never be routed through the NLI classifier tier. The classifier has no trained target for these dimensions, and forcing them through it produces noise dressed up as signal. Those rubrics go straight to the judge, bypassing the cost savings the cascade offers elsewhere, because there is no cheaper proxy that approximates them honestly.

What the deterministic and classifier tier checks

Deterministic checks catch the structural failures that account for a large share of harness regressions, and they do it cheaply enough to run on every PR with no path-scoping needed to keep costs in check. A compiled JSON schema validator operates in microseconds per record. It catches malformed outputs, missing required fields, tool calls referencing deprecated parameter names, and response shapes that would otherwise cause a downstream parser to fail silently in production.

Tool schema drift is a failure category of its own, separate from prompt drift. Prompt drift changes the instructions a model gets, and that changes how the model behaves. Tool drift changes what tools are available, how their schemas are defined, what permissions they carry, or what results they return, and a third-party API can change its response shape on its own, independent of the agent's own codebase. LLM tracers that operate at the application layer cannot see the API execution underneath, so the schema check at this tier is often the only place that catches the drift before it reaches a user.

NLI-backed classifiers extend this coverage into semantic territory without the cost of a frontier model call. Faithfulness, claim support, RAG faithfulness, and factual consistency can all be scored by a classifier trained for natural language inference, and these checks act as the bridge between pure structural validation and the full judge sweep reserved for Tier 2. A lightweight NLI model can run on CPU and clear most semantic checks before any example needs a frontier judge's reasoning.

Measuring determinism directly means running the same input against the agent multiple times and comparing Jaccard similarity across normalized final answers and tool-call sets, which flags agents that produce inconsistent output for identical input. Low determinism usually traces back to underspecified prompts, loosely written tool descriptions, or a model-setting change that went unnoticed, all failures in the harness.

Path-scoped triggering is what keeps even this cheap tier from turning into noise. Running evals only when prompt files, test cases, or model configuration actually change, rather than on every commit, is the specific mechanism that keeps Tier 1 checks relevant. Promptfoo's own GitHub Actions workflow illustrates the pattern directly: it scopes its trigger to prompts/** and promptfooconfig.yaml, so the eval job does not run at all when a commit touches unrelated files.

Scoping the LLM judge to where it earns its cost

The frontier judge tier earns its cost only when it is scoped narrowly and read through a statistical gate calibrated to separate signal from judge variance. Three rubrics belong here specifically because no cheaper proxy approximates them: groundedness, answer refusal, and function-calling accuracy. Each requires the kind of reasoning only a frontier model provides, and routing them through an NLI classifier would produce the same noise problem that applies to helpfulness and tone.

The judge tier runs on two separate cadences. On every PR, it scores only the examples the classifier flagged as low-confidence. On a nightly schedule, a separate cron workflow runs the full suite against the complete versioned dataset across every route, independent of any particular pull request. This nightly sweep updates the stored baseline when the main branch has genuinely shifted and posts the day's results back into the observability system, so the PR gate's floor stays calibrated against real variance instead of a number frozen at some arbitrary point in the past. Slow drift, the kind that never trips a floor threshold because no single day crosses it, is what this nightly cadence is built to surface.

The statistical gate turns the judge's score into a trustworthy decision. Continuous rubrics get a Welch's t-test; binary metrics get a z-test; tail latency gets evaluated at p95. The threshold requires two conditions together: statistical significance at p < 0.05 and a minimum effect floor. Running the two checks in tandem matters because each catches something the other misses. The floor check catches a catastrophic drop immediately, even in a case where the sample is too small to reach formal significance. The delta check, run through Welch's t-test, catches the slow erosion that never drops below the floor on any single PR but builds up across weeks of small changes. Sample size interacts with both: a small enough dataset produces a confidence interval wide enough that a real quality drop sits inside the noise band, and the gate fires on nothing in particular, training engineers to stop trusting it.

A CI policy becomes enforceable when a statistical verdict is translated into exit codes. Exit 0 signals success, exit 1 signals a test or assertion failure, exit 2 signals interrupted execution, and exit 3 signals an internal error in the eval tool itself. Exit 6 and exit 7 vary by tool and aren't standardized across the ecosystem; they typically cover API errors or timeouts, so a workflow that branches on these two codes needs to check its specific eval tool's documentation for their meaning. The workflow reads these codes to decide whether to block the merge outright, post a warning comment and let the author proceed, or retry automatically on what looks like a transient failure.

Building the GitHub Actions YAML that wires all three tiers together

A production-grade eval pipeline is a set of coordinated workflows, each with its own trigger matched to the tier it serves, not a single YAML file trying to do all three jobs at once. Collapsing the cascade into one job recreates the slow, expensive gate the entire tiered design exists to avoid.

The PR-facing workflow handles Tier 1 and the scoped portion of Tier 2. Its trigger fires on pull_request events but filters on paths, limited to prompt files, eval configuration, and test cases, so unrelated changes to the codebase never invoke the job. Deterministic checks run as the first step, and when they pass cleanly the judge step either gets skipped entirely or scoped down to only the examples flagged by the classifier. A quality gate step follows, parsing the results JSON and checking the failure count against a threshold; promptfoo's own pattern uses jq to pull the failure count out of the results file and exits non-zero when it exceeds the configured limit. Artifacts upload unconditionally, with if: always() set on that step, so a failed run still leaves its output available for debugging afterward.

The nightly workflow handles the full Tier 2 sweep, running on a schedule cron trigger against the complete versioned dataset. A red team or security scan is a natural companion job inside this same nightly workflow: adversarial probes run against the current main branch on the same daily cadence, distinct from the quality eval but sharing the same trigger and infrastructure. The baseline update logic lives here too. If the main branch's scores have shifted within expected variance, the stored baseline updates, but if they have dropped past the effect floor, the workflow opens an issue or pages the on-call engineer directly.

Teams already running on the Azure Foundry stack have a first-party option: microsoft/ai-agent-evals@v3-beta accepts agent IDs in agent-name:version format, evaluates multiple agents in a single run, and returns statistical test results directly, folding a meaningful portion of the Tier 2 and gate logic into a single maintained action.

Secrets management across all of these workflows follows one rule without exception: LLM provider API keys live in repository or organization secrets, never written into the YAML body itself, and the workflow references them as ${{ secrets.OPENAI_API_KEY }}. A missing secret does not throw a workflow-level error. It causes a silent failure at the eval step itself. Both secrets need to be confirmed present before the first push, not discovered missing after a confusing green checkmark on a run that did nothing.

What replay-based validation adds beyond the CI gate

Even a well-built cascade is bounded by its dataset, and no curated dataset covers the full distribution of production inputs, tool states, and intermediate agent states that cause real regressions in the field. Replay against recorded production traces covers exactly the ground the golden dataset cannot reach, because it tests against what actually happened, not what the dataset's authors anticipated.

A replay record captures what the agent sent to each tool, what it got back, the decisions any sub-agents made along the way, and the intermediate state carried between steps, the minimal metadata needed to reconstruct a run deterministically after the fact. When the underlying agent code is unchanged, any divergence that appears during replay points to a real correctness issue. That divergence might be a control-flow fork, a prompt that drifted without anyone updating the eval dataset, or a dependency the tracing never captured.

Agents that integrate planning, memory, reflection, and tool use are particularly exposed to a failure pattern replay is suited to catching: a single root-cause error early in a trajectory cascades through every decision that follows it, compounding as the run proceeds. Synthetic golden tasks rarely reproduce this kind of cascading failure, because people build them around isolated test cases, not the long, stateful trajectories real production traffic generates. Recorded production traces capture it as a matter of course.

Replay functions as a debugging discipline as much as a testing step. When an agent fails, replaying the recorded trace shows which harness component, the prompt, the tool schema, the workflow state, or the memory layer, let the wrong state propagate, rather than leaving an engineer to re-run the agent with a tweaked prompt and hope the original behavior reproduces on command. Knowing that a run failed is not, by itself, actionable information. Knowing that the failure originated in tool schema drift at step three of a six-step trajectory is actionable, and replay is the mechanism that preserves the evidence an engineer needs to make an attribution that specific.

The observability instrumentation the pipeline depends on

Step-level tracing underneath it is what makes all of the preceding architecture function. Statistical gating, replay, and root cause attribution all depend on evidence pulled from real agent runs, so that tracing substrate is what lets the pipeline operate on what the agent actually did.

The OpenTelemetry GenAI specification provides the vendor-neutral foundation this depends on, organized across six layers: client spans, agent and workflow spans, MCP conventions, events, metrics, and provider-specific attributes. A common schema across these layers means eval tooling and observability backends can consume the same trace data without either side building a custom adapter to translate between them.

Two histogram metrics function as an effective floor beneath everything else in this pipeline. gen_ai.client.operation.duration, recorded in seconds, carries a Recommended requirement level under the OTel GenAI spec and is the metric that makes p95 tail-latency gating possible. gen_ai.client.token.usage, broken out by input and output tokens, is what makes cost tracking and prompt-bloat detection possible. Without these two metrics captured consistently, a team has no rigorous way to reason about cost or speed at all, and both the nightly baseline update and any canary rollback decision depend on having them in place before the question is ever asked.

Sources

  1. How to run an evaluation in GitHub Action - Microsoft Foundry

More in Online & Offline Eval