Skip to content

12/7/2026 · Clevertek Team

Cloud GPU for AI Inference: Cost and Latency Tradeoffs

How to weigh on-demand cloud GPU against reserved capacity for AI inference, and where latency and cost really diverge for production workloads.

Running AI inference in production is a different problem from training a model in a notebook. Training happens once and tolerates batch delays. Inference is a live service, and every millisecond of added latency shows up in user experience or downstream automation. Cloud based GPU makes inference accessible without buying hardware, but the cost and latency profile depends entirely on how you provision it.

The two cost levers: on-demand vs reserved

On-demand GPU instances let you spin capacity up and down with no commitment. That is ideal for spiky or experimental traffic, but the per-hour rate is the highest available. Reserved or committed capacity lowers the effective hourly cost substantially, at the price of paying for capacity you may not always use.

A simple way to frame the decision:

  • If your inference load is steady, reserved capacity almost always wins on unit cost.
  • If your load is bursty, on-demand avoids paying for idle GPUs but raises blended cost per request.
  • If you are piloting, on-demand keeps you liquid until the traffic pattern is real.

A third shape sits between those and neither extreme handles it well: a steady baseline with a sharp, predictable peak. That pattern wants both. Run the baseline on committed capacity, let on-demand absorb the peak, then check each quarter whether the peak has grown enough to justify committing more of it.

Where latency actually comes from

GPU compute time is only one slice of inference latency. The rest is network round trip, queueing, and cold starts. A model sitting idle on a reserved instance answers in milliseconds. An on-demand instance that has to initialize a runtime or pull a large model into memory can add noticeable lag on the first request.

It helps to break total latency into its parts before optimising anything:

  • Client to endpoint. Network distance between the user and the region serving them. Geography fixes this, and a faster card does not.
  • Queue time. How long the request waited for a free execution slot. That is a capacity and concurrency question rather than a compute question.
  • Cold start. Runtime initialisation plus weights loaded into memory. Present only when the worker was not already warm.
  • Prefill and decode. The model reading the prompt, then generating output token by token.
  • Response transfer. Streaming the result back, which matters more for long outputs than short ones.

Only two of those five respond to a bigger GPU. Spending on accelerators when the bottleneck sits in queue time or geography is the most common way inference budgets get wasted.

Placing inference close to users matters. A cloud based GPU region far from your request origin adds unavoidable network latency that no amount of GPU power removes. Multi-region or edge-adjacent deployment is the lever for that, not a bigger card.

Cold starts are a product problem, not only a performance one

A cold start is invisible in a benchmark that runs a thousand requests back to back, and very visible to the first user after a quiet period. Consumer services feel it at the start of the working day. Internal tools feel it after a lunch break. Both arrive as complaints about a system that measured well in testing.

Three practical mitigations, in rising order of cost:

  1. Keep a minimum number of warm workers. The simplest fix and the most expensive, because you pay for idle capacity around the clock. It is the right answer where a customer-facing latency budget exists.
  2. Snapshot the loaded state. Where the runtime supports it, a pre-loaded snapshot shortens the initialisation window considerably against a cold container start.
  3. Route the first request elsewhere. Accept the cold start on a background job or a non-interactive path, and keep the interactive path warm.

Scale-to-zero is attractive on a cost graph and hostile on a latency graph. Choose it deliberately for workloads where nobody is waiting, and avoid it where somebody is.

Batch size and throughput tradeoffs

Small batch sizes reduce per-request latency but waste GPU parallelism. Larger batches raise throughput and lower cost per inference, at the cost of higher tail latency. Tuning batching policy is often a bigger cost lever than the instance type itself.

The reason is utilisation. A GPU executing one request at a time leaves most of its arithmetic units idle. Batching fills them, so cost per inference falls without any change to hardware. The price of filling them is that the last request admitted to a batch waits for the rest of it.

That trade suits batch scoring, document processing and offline enrichment. It rarely suits an interactive assistant where a person is watching the response appear. The usual resolution is two serving paths: a latency-optimised endpoint for interactive traffic with small batches, and a throughput-optimised endpoint for bulk work with large ones.

Quantisation changes both sides of the equation

Reducing weight precision lowers the memory footprint and speeds up decode, which usually reduces cost per inference. It also changes output quality, and that change is not uniform across tasks.

The honest position is that quantisation is an empirical decision. Summarisation and classification often survive aggressive quantisation with no measurable quality drop. Code generation and reasoning-heavy tasks are frequently more sensitive. Decide with your own evaluation set at each precision level rather than assuming a general rule applies to your workload.

What quantisation reliably does not fix is a latency problem caused by network distance or queueing. It shortens the compute slice, which may be a minority of the total.

Measure cost per thousand inferences

Instance pricing is an input, not an outcome. What the business cares about is what one thousand inferences cost, and that number moves with batching efficiency, quantisation, cache hit rate and how much of the fleet sits idle.

Track these alongside each other:

  • Cost per thousand inferences, split by endpoint and by model.
  • p50 and p95 latency, so a cost improvement that damaged the tail stays visible.
  • GPU utilisation, as the share of provisioned capacity actually doing work.
  • Cache hit rate, where prompt or response caching applies.

When cost per thousand inferences rises while utilisation falls, the fleet is oversized for current traffic. When it rises while utilisation holds, the workload itself changed. Those two situations call for opposite responses, and a single monthly cloud bill cannot tell them apart.

A practical starting point

Start on-demand to learn your real traffic shape. Once diurnal patterns are clear, move the baseline steady-state load to reserved capacity and keep on-demand only for peaks. Measure cost per thousand inferences, not cost per hour, and let that number drive the split.

Two decisions are worth making before the workload reaches production, because both get harder to change afterwards.

Where the endpoint lives. Region placement is a latency decision with a data residency dimension attached. Where the workload processes personal data, the region is not free to choose on latency grounds alone, and the residency question should be settled before the first benchmark rather than after the first release.

Who operates the serving layer. Autoscaling policy, model versioning, rollback and capacity planning are ongoing work. At small scale the team that built the model can carry them. That stops being true once several teams serve several models on shared infrastructure. Deciding early whether that layer is owned internally or handed to a managed platform avoids a disruptive migration later.

Neither decision demands a large upfront commitment. Both are considerably cheaper to make before the traffic arrives than after it.

Ready to modernise your network, cloud and communications?

Talk to Clevertek about a solution scoped to your enterprise — no obligation.

Talk to us