All notes

Engineering

Some things we've learned scaling LLM inference in production

A loose stream of paper documents funnelled through a steel machine and sorted into separate lanes, one of them lit green.

One of the platforms we build and run processes millions of pages of documents every month, using multiple AI agents to automate the workflows around them. It uses a lot of LLM calls.

Most of the volume arrives in a concentrated window and needs to be processed before the review team starts its shift a few hours later. If the pipeline runs late, people are left waiting. Equally, when they are actively using the AI tooling it needs to be fast and responsive. And the AI cannot make dumb mistakes, because users will then second guess everything.

There are naturally a few things fighting against each other:

  • Cost.
  • Model intelligence.
  • Latency.
  • Consistency.
  • Throughput.
  • Provider capacity.
  • Context size.
  • Security and residency.

You could route everything to the best model at full speed and job is done. You will also likely end up in the news with a record token bill.

A few things we've found useful.

1. Decide what each workload actually cares about

If a human is waiting, latency matters. If it is batch work due in three hours, nobody cares if an individual call takes 5 or 20 seconds. They care whether the whole lot finishes on time, at the required quality, and what it costs. Different problem, different optimisation.

2. Put a proper model router in front of the agents

Different steps need different levels of intelligence, and therefore have very different economics. We use routing and guards around sub-agents so the expensive models are available where they actually matter, while high-volume simpler work gets pushed towards cheaper models. The important bit is making that a platform decision rather than trusting every agent or developer to pick sensibly.

3. Batch is where you can go full nerd

Batch gives you much more room to optimise. You can tolerate higher latency per call, use cheaper or open-weight models for high-token and lower-intelligence steps, run multiple specialist agents, add validation passes and retry selectively. One model call doesn't need to be perfect if the pipeline around it is designed properly.

4. Scale against the queue, not average traffic

Average utilisation can look amazing while your queue is on fire for a couple of hours a day. For bursty workloads, the real question is less "how much traffic do we get?" and more "how much work has to be finished by when?" That means looking at queue depth, queue age and throughput, then scaling GPU workers or provider concurrency around the backlog and deadline rather than just incoming request count. For the really high-volume parts, open-weight models on elastic GPU capacity can change the economics considerably.

5. Assume providers will eventually ruin your evening

Quotas get hit. Capacity disappears. Models get slower. APIs have bad days. A routing layer gives you somewhere to manage quotas, failover and provider-specific weirdness without teaching every workflow about it.

6. Measure cost per useful thing

Tokens are useful for engineers. "How much does it cost to process one document or complete one workflow?" is usually the more useful question. Add this early and track it by task and model. Otherwise finding the expensive bits later becomes surprisingly painful. That also makes it fairly obvious where you're paying Range Rover prices to drive to the shops.

7. Don't make the model do the same work twice

LLM calls are expensive enough the first time. Use the boring stuff properly: checkpoint stages, make retries idempotent, cache repeated inference, and resume from the last successful step rather than restarting an entire workflow because one call failed. At volume, those basics make a surprisingly large difference to both reliability and cost.

8. Make evals business-led

Having another model available is not the same thing as having a valid fallback. A model looking good on a generic benchmark doesn't mean it is good enough for your workflow. Those late nights spent prompt hacking often mean the old model still beats the new one. So before any model goes into the routing pool, test it against business-led evals that reflect the actual mistakes users care about.

A lot of this isn't really new computer science. It is queueing, routing, capacity planning and making sensible trade-offs between intelligence, latency and cost. LLMs just make the bill more interesting.

Working on something like this?

If any of the above sounds familiar, we have probably run into it before.

Start a conversation