← Notes

How to compare LLM API prices properly

Per-million-token prices look comparable and are not. Output usually costs several times input, tokenisers disagree about how long your text is, and caching changes the bill more than the headline rate.

Every provider publishes a price per million tokens. Two models at "$3 / $15" and "$2.50 / $10" look like a straightforward comparison, and almost nobody who makes that comparison ends up with the bill they expected.

Here is what the headline number leaves out.

Input and output are different products

Almost universally, output tokens cost several times more than input tokens — often three to five times. The two numbers you see quoted are input / output, and which one dominates your bill depends entirely on the shape of your workload.

A summariser feeds in ten thousand tokens and writes back three hundred. It is an input-heavy workload, and the input price is what matters. A generator that takes a one-line prompt and writes a thousand-word article is output-heavy, and the input price is nearly irrelevant.

So the first thing to work out is not "which model is cheaper" but what your input-to-output ratio actually is. Log it from a real sample of requests. Guessing is how people end up surprised.

A token is not a word, and not the same everywhere

Pricing is per token, and a token is roughly three-quarters of an English word. Roughly. Two things break that rule of thumb badly:

Non-English text costs more. Tokenisers are trained predominantly on English. The same sentence in Turkish, Japanese or Arabic often takes noticeably more tokens than its English equivalent — sometimes twice as many for agglutinative or non-Latin-script languages. If your product serves a non-English market, the English-based estimate understates your bill.

Structured text costs more than prose. JSON, code, tables and anything with lots of punctuation tokenise less efficiently. A prompt full of JSON schema is more expensive per useful word than the same instruction in a sentence.

Different providers also use different tokenisers, so the identical prompt is a different number of tokens at each. You cannot compare two prices without comparing them against the same text run through each provider's own counter.

The things that move the bill more than the rate

Prompt caching. If you send the same long system prompt or document on every call, most providers will cache it at a large discount — frequently a 90% reduction on the cached portion. For a chatbot with a big system prompt this is often the single biggest lever available, and it can make a nominally more expensive model cheaper in practice. It usually requires you to structure the prompt so the stable part comes first.

Batch processing. Work that does not need an answer within seconds can often go through a batch endpoint at around half price. Overnight classification, backfills, evaluation runs — all good candidates.

Reasoning tokens. Models that think before answering bill those intermediate tokens as output, and you never see them. A model quoted at a low output rate can cost several times more per answer than one quoted higher, simply because it generates far more tokens to get there. Compare cost per completed task, not cost per token.

Retries and failures. Timeouts, rate-limit backoffs and validation failures that force a second call are real cost. A cheaper model that needs two attempts to produce valid JSON is not cheaper.

A comparison that actually tells you something

  1. Take fifty real requests from your application — not synthetic ones.
  2. Run them through each candidate model.
  3. Record input tokens, output tokens (including reasoning tokens if the API reports them), and whether the result was usable.
  4. Compute cost per successful result.
  5. Then apply caching and batching to whichever workload allows it, and recompute.

That number is comparable. The published rate is not.

Two more things worth checking before you commit: the rate limits attached to your tier, which can make a cheap model unusable at your volume, and whether prices are per-region. Both have a way of appearing after you have built.

Where AIonRadar helps

AIonRadar tracks AI model and API releases with links back to the original sources, so the numbers you are comparing are the ones the provider actually published rather than a figure copied from a blog post six months ago.

It also has the two calculators this article keeps circling: one for estimating API usage cost from your own token volumes, and one for GPU and VRAM requirements if you are weighing self-hosting against an API. The LLM API Selector matches routes to a described workload, which is a faster starting point than reading nine pricing pages.

It is free on the web, and the iPhone and iPad app has no account, no advertising and no in-app purchases.

Keep reading