What is running
- Inbound triage
- Reads everything arriving, classifies it, routes it, drafts a reply. A person sends. Highest value of the five by a distance.
- Scheduling
- Holds slots, handles rescheduling, chases confirmations. Boring, reliable, and nobody has thought about it in months.
- Document intake
- Extraction and classification into the system of record. Gated above a threshold. The one with the most engineering behind it.
- Exception summarisation
- When something needs a person, it arrives with the context assembled. Not a decision-maker, a preparer.
- Weekly reporting
- Pulls from live data and drafts the narrative. Saves an afternoon a week and is checked in ten minutes.
What we removed
Two workflows were switched off, and the reasons generalise.
The first was an agent that decided which work to prioritise. It worked, in the sense that its rankings were defensible. It was removed because nobody trusted it and everybody quietly re-sorted the list, so it added a step and changed nothing. The lesson: automating a judgement that people will redo anyway is worse than not automating it.
The second was a fully autonomous responder on a narrow class of routine enquiries. Accuracy was fine. It was removed because the failures, though rare, were embarrassing in a way that cost more than the time saved. We moved it behind the same draft-and-approve pattern as everything else, where it still runs.
Both removals came from the same root: we automated the commit rather than the work. The pattern that survives everywhere is that the system does the work and a person commits it.
What surprised us
- The operational scaffolding was most of the effort, and the agents were the small part. Queues, retries, idempotency, the decision store and the review interface took roughly four times the model work.
- The review interface mattered more than the accuracy. When we made the evidence visible next to the claim, review time fell by about two thirds and the correction rate went up, because people were actually reading it.
- Override data turned out to be the most valuable thing we collected. It is a labelled eval set arriving continuously, and it is what tells us when something has drifted well before any aggregate metric moves.
- Nobody wanted fewer people. The team took on work they had been deferring for a year. That is what the saving actually bought, and it is a better thing to promise a client than headcount.
What it changed about how we scope
Running it ourselves changed the advice we give, in three concrete ways.
- We now insist on the decision store before the first workflow ships. We did not, once, and spent a fortnight reconstructing why something had happened from logs that were never designed for it.
- We scope the review interface as a first-class deliverable with its own budget, rather than as a screen someone adds at the end.
- We start every engagement in draft-and-approve and agree the ratchet in writing on day one. Every category we have since automated came off that evidence, and the conversations were short because the numbers already existed.
The uncomfortable version of this: most of what makes an AI system work in production has very little to do with AI. That is not a marketing message, but it is what a year of running our own has taught us, and it is why the advice comes from operating rather than from a methodology.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
