Skip to content
ENع
Schedule call
All guides
Models and prompting8 min read

Choosing a model for a task

There is no best model. There is a best model for one step of one workflow at one point on the cost curve, and it changes about twice a year. Picking per project is what makes the decision feel hard.

Three axes, and they genuinely trade

Every model choice moves along three axes at once, and improving one usually costs you another. Naming them stops the argument going in circles.

Reasoning depth
Can it hold a multi-step chain without losing the thread? This is what separates a model that classifies an invoice from one that reconciles a disputed one.
Latency
Time to first token and time to complete. A human waiting on a screen tolerates about two seconds. A queue worker tolerates minutes.
Cost per call
Input and output priced separately, and output is several times dearer. A step that reads a lot and writes a little is much cheaper than the token count suggests.

A fourth axis matters in the Gulf specifically: where the weights can run. A model that cannot be deployed in your tenancy is not a candidate for a workload that cannot leave it, whatever it scores.

Decide per step, not per project

A single workflow usually spans three or four model calls with completely different requirements. Treating them as one decision forces you to buy the most expensive model for every step.

An invoice intake pipeline we built runs four calls. Extraction from the scan is a wide, shallow task on a fast cheap model. Classification against the chart of accounts is a lookup with a small model and a good retrieval layer. Reconciliation of a mismatch is the only step that needs real reasoning depth. The summary written back to the operator is a cheap model again.

Putting the reasoning model on all four would have cost roughly nine times as much for no measurable accuracy gain on three of them. We know that because we measured it, which is the next section.

The test that settles it

Benchmarks answer a question you do not have. Run this instead, and it takes an afternoon.

  1. Pull fifty real cases from your own history, decided, with the right answer known. Include the hard ones; they are the whole point.
  2. Write the prompt once. Do not tune it per model, or you are measuring your tuning.
  3. Run all fifty through two or three candidate models. Record accuracy, p95 latency and total spend.
  4. Split the accuracy by whether the case was routine or an exception. This is where models separate, and where an average hides everything.
  5. Pick the cheapest model that clears your bar on the exceptions. On the routine cases almost everything clears.

If two models are within a point of each other on your own data, they are the same model for your purposes. Take the cheaper one and spend the difference on evals.

When the answer is no model

A surprising share of steps that arrive framed as AI problems are not. Saying so early is cheaper for everyone.

  1. The rule is stable and writable. If a person can state the rule in a sentence and it has not changed in two years, write the rule. It will be faster, free and testable.
  2. The input is already structured. If the data arrives as a well-formed record, you want a validator, not a language model.
  3. The cost of a wrong answer exceeds the cost of a human doing it. At low volume with high stakes, automation is the wrong shape regardless of accuracy.
  4. The real problem is a missing field. We have twice found that cleaning one column upstream removed the need for the model entirely.

Plan for the swap

Whatever you pick will be superseded. The engineering that matters is the part that makes replacing it a config change rather than a project.

  1. Keep the model behind an interface that takes your types, not the provider’s. One adapter per provider, and the workflow never imports an SDK.
  2. Keep prompts in version control as data, not embedded in code paths.
  3. Keep the eval set. It is the only thing that tells you whether the new model is actually better on your work, and it is what makes the swap a one-day job.
  4. Record which model decided what, on every decision. When behaviour changes after a swap you will want to know which calls were affected.

Done that way, the model is the cheapest component to replace. Done the other way, it is the project.

Want us to run this with you?

The Audit is this method pointed at your systems, with a costed build plan at the end of it.

Schedule call
Tell us the number you want to move.Schedule call