What a case needs
- The input as received
- The original document, email or record, not a cleaned version. If production sees a bad scan, so should the eval.
- The decision that was made
- The actual outcome, from the system of record.
- Who decided, and when
- For provenance, and because you will find that two people decided differently on similar cases.
- Tags
- At minimum whether it was routine or an exception, and what kind. This is what makes the set useful rather than merely large.
Two hundred well-chosen cases beat ten thousand sampled at random, because the ten thousand will be 95% routine and tell you nothing you did not know.
Sample for the hard cases deliberately
If you sample uniformly, your eval set has the same distribution as your work, which means the score is dominated by cases everything gets right. A model that is 96% accurate overall and 40% accurate on exceptions is a model that will fail in production, and a uniform sample will report 96%.
- Take a hundred routine cases, sampled randomly. This is your floor; everything should pass.
- Take every exception type you can name, ten to twenty of each. Multi-page documents, disputed amounts, unknown vendors, foreign currency, whatever your operators complain about.
- Add the cases your team argued about. Ask them; they remember. These are where the specification is genuinely unclear and they are the most valuable rows in the set.
- Add the ones that were decided wrongly and later corrected. The correction is the label.
- Tag every one. You will score by tag, not in aggregate.
Score exceptions separately and hold them to their own bar. An overall number that mixes them is the metric that lets a broken system look fine.
Handling disagreement
You will find cases where two experienced people decided differently. This is not a problem with your data collection; it is the most useful thing the exercise produces.
- Do not average them and do not pick one arbitrarily. Take them to the person who owns the policy.
- Sometimes the answer is that both are acceptable. Then the eval accepts either, and you have learned the specification is a range.
- Sometimes the disagreement reveals the rule was never written down. Write it down. This is often worth more to the client than the model.
- If a case cannot be resolved, exclude it and record why. An eval set with unresolved cases produces a score nobody trusts.
Keep it honest over time
- Hold a slice back. If every case is used to tune prompts, your score measures your tuning.
- Refresh quarterly from recent decisions. Input distributions move, and a two-year-old set slowly stops describing the work.
- Version it alongside the code, so a score is always attributable to a specific set.
- Hand it to the client’s team, and mean it. An eval set they cannot run themselves is a dependency on you, and the whole point is that it should not be.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
