Alert on behaviour, not just on health
Uptime, latency and error rate will all look fine during your worst incident. These are the signals that move.
- Override rate by category
- The best single indicator you have. A step change means the system and its operators have started disagreeing, which is what going wrong looks like.
- Confidence distribution
- Alert on the shape, not the mean. A bimodal distribution flattening out is drift arriving.
- Output distribution
- If 5% of invoices were normally routed to one code and today it is 30%, that is an incident whatever the accuracy metric says.
- Abstention rate
- A model suddenly declining more is telling you something changed upstream.
- Cost per case
- A jump usually means a retry loop or a prompt that grew. Both are bugs.
- Exception queue age
- Because the failure mode of a gated system is a queue nobody is clearing.
The runbook, written for somebody else
The person on call did not build this. Write for them, and test it by having them use it.
- For each alert: what it means in one sentence, the three most likely causes in order, and the first thing to check.
- How to stop it, including degraded modes, and who is allowed to.
- How to find the decisions from the last hour and read them. This is the first diagnostic step for nearly every AI incident.
- The provider status page, and how to tell a provider incident from your own. This is the most common ambiguity at 2am.
- Who to wake, and at what threshold, by name.
A runbook is tested by handing it to someone who has never seen the system and asking them to work an alert. Anywhere they stop is a gap, not a training problem.
Incidents that are specific to this
- Provider degradation
- Not down, just slower or subtly worse. Your latency alert may not fire. Watch quality signals and have a fallback model configured before you need it.
- Silent model change
- A floating alias moved under you. This is why pinning versions matters. Symptom: behaviour changed with no deployment.
- Input drift
- A supplier changed their invoice template. Accuracy falls only for them, so aggregates hide it. Per-source metrics catch it.
- Prompt injection
- Content instructing the model. Symptom: tool calls that make no sense for the case. Alert on unusual tool sequences.
- Retry storm
- A transient failure amplified into sustained load and spend. Symptom: cost per case jumping while volume is flat.
Handing it over properly
Our engagements end with the client’s team running this without us. That is a deliverable, and it either happens deliberately or it does not happen.
- They hold the pager from the first week of production, with us alongside. Shadowing after the fact does not transfer anything.
- They run the eval set, on their machines, in their CI, and can read the result.
- They make the first prompt change, with us reviewing rather than typing.
- They work a real incident while we are still there, so the runbook is proven rather than believed.
- The last thing we do is a session on what we would worry about next. It is usually the most valuable hour of the engagement.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
