Skip to content
ENع
Schedule call
All insights
Method7 min read

Evals your operators can run themselves

A vendor benchmark tells you how a model performs on somebody else’s work. It has almost no predictive power for yours, and the gap is widest exactly where it matters: the difficult cases your business actually loses money on.

Why the public number does not transfer

  1. It is measured on clean, general, English-language inputs. Yours are scanned, domain-specific and frequently bilingual.
  2. It is dominated by routine cases, because that is what a broad benchmark contains. Your value is in the tail.
  3. It cannot express your cost asymmetry. Missing a fraudulent transaction and flagging a good one are both errors and are not the same error.
  4. It is a target that has been optimised against. Models are tuned toward published benchmarks, which is precisely what makes them poor predictors of unseen work.

We have watched a model that led every public leaderboard score 41% on a client’s exception cases, while a cheaper one scored 78%. The exceptions were bilingual and the leader was not. No benchmark would have told them that.

The set already exists

Every case a client has already decided is an input with a known correct answer, labelled by the people whose judgement the system is meant to reproduce. It is better than anything purchasable and it is usually sitting unqueried in a database.

Two hundred well-chosen cases beat ten thousand random ones. A hundred routine, then ten to twenty of every exception type the operators can name, then the ones the team argued about, then the ones decided wrongly and later corrected. Tag every one, and score by tag.

Scoring in aggregate is what lets a broken system look fine. A model at 96% overall and 40% on exceptions will fail in production and report 96%.

The part that matters: they run it

An eval set the client cannot run is a dependency on the supplier, which defeats the purpose. Ours are handed over in the same shape we use them.

  1. In their repository, in their CI, on their credentials.
  2. One command, with output a non-specialist can read: score by category, delta against the baseline, cost per case.
  3. Documented well enough that their team adds a case without asking us. Adding cases is the habit that keeps it alive.
  4. Wired to a gate, so a change that regresses a category cannot merge.

What it changes

A client with their own eval set can change models, change suppliers, change prompts and change us, safely. That is a slightly uncomfortable thing for a consultancy to build and it is the main reason clients trust the rest of the work.

It also changes the conversation. Arguments about whether the system is good enough become arguments about a number, and arguments about a number can be settled.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call