← Notes

Hosted API or local inference: choosing honestly

The break-even point is further away than enthusiasts claim and closer than vendors imply. It depends on your volume, your privacy constraints, and whether you have counted your own time.

Two camps give confident answers to this. One says APIs are a tax and you should own your hardware. The other says running your own inference is a hobby project pretending to be infrastructure. Both are right about specific situations and wrong as general advice.

The arithmetic, done properly

API cost is straightforward: tokens in, tokens out, multiplied by published rates. Apply caching and batch discounts where your workload allows them, and you have a monthly figure that scales linearly with usage.

Local cost is not one number:

  • Hardware. A GPU with enough memory for a useful model, or a Mac with enough unified memory. This is the number people quote.
  • Electricity. A card under sustained load draws real power. At European rates, continuous inference on a large card is a meaningful monthly line.
  • The machine around it. Local inference on your laptop competes with your actual work for memory and thermal headroom.
  • Your time. Setup, model updates, quantisation choices, runtime upgrades that break things, debugging why throughput dropped. This is the cost that never appears in comparisons and is usually the largest one.

The honest comparison is total cost of ownership over the period you will actually use it, against the API bill for your real volume — not for the volume you imagine at scale.

Where the break-even usually sits

For most individuals and small teams, an API is cheaper than hardware until usage is both high and steady. Bursty usage is the worst case for local: the hardware costs the same whether it runs for eight hours a day or eight minutes.

The rough shape:

  • Occasional use — a few hundred to a few thousand calls a month. API, comfortably. The hardware would take years to pay back.
  • Steady high volume — continuous classification, bulk processing, an internal tool many people hit all day. Local starts winning, especially with a model small enough to run on modest hardware.
  • Experimental and exploratory — API, because you will want to try five models this month and buying hardware commits you to one size class.

The reasons that are not about cost

These override the arithmetic entirely, in both directions.

Data that cannot leave. Client confidentiality, medical or legal material, a contract that forbids third-party processing. If the data cannot go to a vendor, the comparison is over. Local is the only option, at whatever it costs.

Reliability requirements. An API is a dependency with its own outages, rate limits and deprecation schedule. A model on your own disk does not get retired. If a product must keep working through a provider's bad day, or must produce identical output in two years, local has an argument that money does not.

Latency to the user. Local inference on the user's own machine removes the network round trip entirely, which for interactive features is a different experience rather than a cheaper one.

Quality ceiling. Against all of the above: the best open-weights models are genuinely good and are not the best models available. If your feature needs frontier capability, local is not a substitute at any price.

The third option people skip

Not everything has to go the same way. A sensible split is common:

  • Bulk, repetitive, privacy-sensitive work runs locally on a small model.
  • Hard, rare, quality-critical requests go to an API.

This costs more engineering effort than either alone, and it is usually the cheapest answer for a product with a mixed workload.

Before you buy anything

Run your real workload against an API for a month. You will discover your actual token volumes, your input-to-output ratio, and how much of your traffic is easy enough for a small model. Those three numbers make the decision for you.

Buying hardware first and finding the workload afterwards is how expensive cards end up idle.

Where AIonRadar helps

This decision needs two numbers you have to calculate rather than look up, and AIonRadar has a calculator for each: API usage cost from your own token volumes, and GPU and VRAM requirements for the model and configuration you are considering. Putting those side by side is the comparison this article describes.

The LLM API Selector narrows the hosted options to routes that fit a described workload, which is faster than reading nine pricing pages, and the news side tracks model and hardware releases with links back to the original sources — so the specifications you plan around are the published ones.

Free on the web, with an iPhone and iPad app that has no account, no advertising and no in-app purchases.

Keep reading