Skip to content
ENع
Schedule call
All guides
Evals and quality7 min read

Regression gates in CI

The moment a prompt can be changed without a score attached, it will be, usually to fix one complaint, and it will quietly break three things nobody was watching. A gate is what turns an eval set from a report into a control.

Two tiers, because the full set is too slow

A few hundred model calls is minutes and real money. Run that on every push and people will start working around it.

On every push
Fifty cases, chosen to cover every category, weighted to exceptions. Under two minutes. Blocks the merge.
Nightly and before release
The full set, every category, with the trend recorded. Blocks the release.
On a model or provider change
The full set plus the permission and safety tests, always, regardless of how small the change looks.

Gating a non-deterministic suite

The same input can produce a different output, so a strict pass or fail on individual cases produces a flaky gate that people learn to rerun until it is green. That is worse than no gate.

  1. Gate on the aggregate score by category, not on individual cases. "Exceptions at or above 85%" is stable in a way "case 41 passes" is not.
  2. Set the bar from the current score, not from an aspiration. The gate exists to stop regression, not to encode a target.
  3. Allow a small tolerance, a point or so, to absorb sampling noise, and set temperature to zero where the provider supports it.
  4. Fail loudly on a category collapse. A single category dropping twenty points is a real regression even if the average holds, and this is the check that catches most genuine breakage.
  5. Print the delta per category in the CI output. A reviewer should see what moved without opening anything.

Never rerun to get green. If a gate is flaky, the fix is a larger sample or a wider tolerance, decided deliberately, not a second attempt.

What else belongs in the gate

  1. The permission tests from the retrieval guide. They fail rarely and catastrophically, which is exactly what a gate is for.
  2. Cost per case. A prompt change that improves accuracy by a point and triples the bill should be a conscious decision, not a surprise in the invoice.
  3. Latency at p95, for anything a person waits on.
  4. A schema conformance check on every structured output, which catches contract drift immediately.

Make the result reviewable

The gate’s output is read by a human deciding whether to merge. Write it for them.

A good CI comment says what changed, by category, against the baseline, with the cost delta and a link to the failing cases. "Exceptions 71% to 84%, routine unchanged, cost per case up 0.4 fils" is a sentence a reviewer can act on. A green tick is not.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call