Skip to content
ENع
Schedule call
All guides
Models and prompting7 min read

When to fine-tune, and when not to

Fine-tuning is proposed in roughly a third of the scoping calls we take and is the right answer in perhaps one in ten of those. It is not that it does not work. It is that it is usually solving a problem that a prompt, a retrieval layer or a cleaner input would solve faster and reversibly.

What it is actually good at

Three problems where fine-tuning earns its cost, and they have a shape in common: the thing you want is hard to say and easy to show.

Format and voice
You need output in a house style or a rigid structure that takes three hundred words to describe and still drifts. A few hundred examples fix it and shorten every call thereafter.
A closed domain vocabulary
Your field uses terms the base model reads as something else. Common in shipping, insurance and clinical coding. Retrieval helps; tuning helps more.
Latency and cost at volume
A small tuned model matching a large prompted one on your narrow task, at a fraction of the price. This only pays above real volume, and you should have the numbers before you start.

What it will not fix

  1. Facts. A tuned model is not a database and will state your old prices with total confidence. Facts belong in retrieval, where they can be updated.
  2. Reasoning depth. Tuning teaches a model what your answers look like, not how to think further than it could before. If the base model cannot follow the chain, examples of correct conclusions will teach it to guess in the right style.
  3. A messy input. If extraction fails because the scans are poor, tuning learns your error distribution. Fix the input.
  4. An unclear specification. If two of your own experts disagree on the right answer, no amount of training data resolves it, and building the training set will surface the disagreement painfully.

The tell is whether you can write down the rule. If you can, prompt it. If you can only point at examples and say "like these", tuning is on the table.

Try these first, in this order

  1. A better prompt with two hard examples. Costs an hour and settles more cases than expected.
  2. Retrieval over your own material, so the model has the facts rather than being asked to remember them.
  3. Decomposition: split the step into two calls with narrower jobs. Most "the model cannot do this" turns out to be "the model cannot do these three things at once".
  4. A bigger model for that one step only. Often cheaper than a tuning project and available this afternoon.
  5. Then tuning, with the eval set you built while doing the four above.

By step five you have an eval set, a clear specification and a baseline, which is exactly what a tuning run needs. Teams that start at step five have none of them and cannot tell whether it worked.

If you do it, plan the upkeep

A tuned model is a fork. It stops moving when the frontier does not, and the gap widens.

  1. Keep the training set in version control with its provenance. You will rebuild it against a new base model within the year.
  2. Keep a prompted baseline running on the same evals. When the next base model beats your tuned one, you want to find out from a dashboard, not a customer.
  3. Budget for re-tuning as an operating cost, not a one-off project cost.
  4. Know your exit. If the tuned model is the only thing that works and it cannot be reproduced, you have swapped a vendor lock for a worse one of your own making.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call