Find out where it goes first
Almost every team we have looked at was wrong about which step cost the most, usually by a factor of five. Measure before you optimise.
- Record input tokens, output tokens, cached tokens and model per call, tagged with the workflow step.
- Group by step, not by day. One step is normally 60% or more of the bill.
- Split input from output. Output is several times dearer, so a chatty step with a short input can dominate a step that reads a whole document.
- Look at the tail. A p99 that is twenty times the median usually means a retry loop nobody has noticed.
Prompt caching, which is most of the win
If your system prompt carries instructions, a schema, examples or retrieved context, you are paying full price to send the same tokens repeatedly. Caching that prefix is typically the single largest reduction available.
- Put everything stable at the front: tools, then instructions, then reference material. Put anything that varies at the very end.
- Mark the boundary at the end of the stable part. Marking it at the end of the whole prompt means every request writes a new entry and none is ever read.
- Check nothing varies inside the prefix. A timestamp, a request id or a reordered JSON key invalidates everything after it, and this is the most common reason caching appears not to work.
- In a conversation, mark the end of the most recent turn as well, so a long thread re-reads its own prefix instead of paying for it again.
A cache read is a fraction of the price of a fresh read, and a write is slightly more than one. If your traffic re-reads a prefix even a few times, it pays immediately.
Routing and ceilings
Routing: send the easy majority to a small model and escalate only what needs depth. You already have the signal for this if you have been storing confidence. In the pipelines we run, between 70% and 90% of cases never need the expensive model.
Ceilings are not an optimisation, they are a safety device, and every public endpoint needs them:
- A hard cap on tokens per conversation or per job, enforced in code, not in a prompt.
- A bound on tool-call rounds within one turn, so a loop cannot run indefinitely.
- Rate limits per caller, held somewhere shared rather than in the memory of one instance.
- A daily spend alarm that pages someone, at a level you would notice but survive.
Every one of those exists because of an incident somebody had. The loop bound in particular: an agent that can call a tool that produces work for itself will do so until something stops it.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
