Skip to content
ENع
Schedule call
All insights
Field note9 min read

Five AI employees, one operations team

Gridwork runs its own operations on the platform we build for clients. Five agent workflows, one small operations team, for about a year now. Some of it worked better than we expected and some of it we have since deleted, which is the more useful half of the story.

What is running

Inbound triage
Reads everything arriving, classifies it, routes it, drafts a reply. A person sends. Highest value of the five by a distance.
Scheduling
Holds slots, handles rescheduling, chases confirmations. Boring, reliable, and nobody has thought about it in months.
Document intake
Extraction and classification into the system of record. Gated above a threshold. The one with the most engineering behind it.
Exception summarisation
When something needs a person, it arrives with the context assembled. Not a decision-maker, a preparer.
Weekly reporting
Pulls from live data and drafts the narrative. Saves an afternoon a week and is checked in ten minutes.

What we removed

Two workflows were switched off, and the reasons generalise.

The first was an agent that decided which work to prioritise. It worked, in the sense that its rankings were defensible. It was removed because nobody trusted it and everybody quietly re-sorted the list, so it added a step and changed nothing. The lesson: automating a judgement that people will redo anyway is worse than not automating it.

The second was a fully autonomous responder on a narrow class of routine enquiries. Accuracy was fine. It was removed because the failures, though rare, were embarrassing in a way that cost more than the time saved. We moved it behind the same draft-and-approve pattern as everything else, where it still runs.

Both removals came from the same root: we automated the commit rather than the work. The pattern that survives everywhere is that the system does the work and a person commits it.

What surprised us

  1. The operational scaffolding was most of the effort, and the agents were the small part. Queues, retries, idempotency, the decision store and the review interface took roughly four times the model work.
  2. The review interface mattered more than the accuracy. When we made the evidence visible next to the claim, review time fell by about two thirds and the correction rate went up, because people were actually reading it.
  3. Override data turned out to be the most valuable thing we collected. It is a labelled eval set arriving continuously, and it is what tells us when something has drifted well before any aggregate metric moves.
  4. Nobody wanted fewer people. The team took on work they had been deferring for a year. That is what the saving actually bought, and it is a better thing to promise a client than headcount.

What it changed about how we scope

Running it ourselves changed the advice we give, in three concrete ways.

  1. We now insist on the decision store before the first workflow ships. We did not, once, and spent a fortnight reconstructing why something had happened from logs that were never designed for it.
  2. We scope the review interface as a first-class deliverable with its own budget, rather than as a screen someone adds at the end.
  3. We start every engagement in draft-and-approve and agree the ratchet in writing on day one. Every category we have since automated came off that evidence, and the conversations were short because the numbers already existed.

The uncomfortable version of this: most of what makes an AI system work in production has very little to do with AI. That is not a marketing message, but it is what a year of running our own has taught us, and it is why the advice comes from operating rather than from a methodology.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call