Two tiers, because the full set is too slow
A few hundred model calls is minutes and real money. Run that on every push and people will start working around it.
- On every push
- Fifty cases, chosen to cover every category, weighted to exceptions. Under two minutes. Blocks the merge.
- Nightly and before release
- The full set, every category, with the trend recorded. Blocks the release.
- On a model or provider change
- The full set plus the permission and safety tests, always, regardless of how small the change looks.
Gating a non-deterministic suite
The same input can produce a different output, so a strict pass or fail on individual cases produces a flaky gate that people learn to rerun until it is green. That is worse than no gate.
- Gate on the aggregate score by category, not on individual cases. "Exceptions at or above 85%" is stable in a way "case 41 passes" is not.
- Set the bar from the current score, not from an aspiration. The gate exists to stop regression, not to encode a target.
- Allow a small tolerance, a point or so, to absorb sampling noise, and set temperature to zero where the provider supports it.
- Fail loudly on a category collapse. A single category dropping twenty points is a real regression even if the average holds, and this is the check that catches most genuine breakage.
- Print the delta per category in the CI output. A reviewer should see what moved without opening anything.
Never rerun to get green. If a gate is flaky, the fix is a larger sample or a wider tolerance, decided deliberately, not a second attempt.
What else belongs in the gate
- The permission tests from the retrieval guide. They fail rarely and catastrophically, which is exactly what a gate is for.
- Cost per case. A prompt change that improves accuracy by a point and triples the bill should be a conscious decision, not a surprise in the invoice.
- Latency at p95, for anything a person waits on.
- A schema conformance check on every structured output, which catches contract drift immediately.
Make the result reviewable
The gate’s output is read by a human deciding whether to merge. Write it for them.
A good CI comment says what changed, by category, against the baseline, with the cost delta and a link to the failing cases. "Exceptions 71% to 84%, routine unchanged, cost per case up 0.4 fils" is a sentence a reviewer can act on. A green tick is not.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
