← Notes

Estimating token cost before you build

A feature that costs half a cent per call is fine until ten thousand people use it daily. Here is the arithmetic to do first, and the four things that make real bills exceed the estimate.

The demo costs nothing. Ten calls while you build, a few cents, and the pricing page never enters your head. Then the feature ships and the bill arrives with a shape nobody modelled.

The estimate takes fifteen minutes and is worth doing before the architecture is fixed, because the cheapest fixes are structural.

The calculation

cost per call = (input tokens × input rate) + (output tokens × output rate)
monthly cost  = cost per call × calls per user per month × users

Two inputs you need to be honest about:

Input tokens is everything you send: system prompt, few-shot examples, retrieved documents, conversation history, and the user's actual message. The user's message is usually the smallest part. A 2,000-token system prompt sent on every call is 2,000 tokens every call, forever.

Output tokens is what comes back — plus, on reasoning models, the thinking tokens you never see and are billed for. Those can be several times the visible answer.

Remember that output typically costs three to five times input, so the ratio matters as much as the totals.

Do it with real numbers, not estimates

Token counts are hard to guess and easy to measure. Run fifty genuine requests through the API and record what it reports. Every provider returns usage in the response.

Guessing is where estimates go wrong by a factor of three, usually because the system prompt turned out to be longer than remembered, or because conversation history grows.

Which is the first of the four things that break the model.

What makes real bills exceed the estimate

1. Conversation history compounds. In a chat feature, each turn resends everything before it. Turn ten sends nine previous exchanges. Cost per turn grows linearly and cost per conversation grows quadratically. A twenty-turn conversation can cost ten times what the first turn suggested.

The fix is structural: truncate history, summarise older turns, or use a sliding window. Decide this before launch.

2. Retries. Timeouts, rate limits, malformed output that fails validation. Every retry is a full-price call. A model that produces valid JSON 90% of the time costs 11% more than its sticker price, not 10% less than a pricier model that always does.

3. Non-English text. Tokenisers are English-centric. The same content in Turkish, Japanese or Arabic can take substantially more tokens — sometimes close to double for non-Latin scripts and agglutinative languages. If you estimated in English and serve a Turkish market, your estimate is low.

4. The users you did not model. Usage is never evenly distributed. A small fraction of users will use the feature ten or twenty times more than the median. Model the mean and the tail, because the tail is where the bill lives.

The levers that actually reduce it

In rough order of impact:

Prompt caching. If a long system prompt or document is identical across calls, most providers cache it at a steep discount — often around 90% off the cached portion. For anything with a big stable prefix this is the single largest saving available, and it requires structuring prompts so the stable part comes first.

Batching. Work that does not need an immediate answer often goes through a batch endpoint at roughly half price. Classification runs, backfills, evaluations.

A smaller model for the easy cases. Most workloads have a distribution of difficulty. Routing the simple 80% to a cheap model and escalating the rest is frequently a large saving for a small amount of routing logic.

Shorter outputs. Ask for what you need. "Answer in one sentence" costs a fraction of an unconstrained answer, and output is the expensive side.

Shorter system prompts. Every token is paid on every call. Prompts accumulate instructions during development that nobody removes afterwards. Re-read yours.

Set a budget before you ship

Two practical safeguards:

  • Spend alerts at the provider, at a threshold you choose deliberately.
  • A per-user rate limit in your own application. Without one, a single script hammering your endpoint is an unbounded bill.

Both take an afternoon and cost far less than the incident that justifies them.

Where AIonRadar helps

AIonRadar has a calculator for API usage cost, which does the arithmetic above from your own token volumes rather than leaving you with a spreadsheet — and a GPU and VRAM calculator for the moment the numbers get large enough that self-hosting enters the conversation.

The LLM API Selector narrows the choice to routes that suit a described workload, and the news side tracks model and pricing announcements with links to the original sources, which matters because rates change and blog posts do not.

Free on the web, with an iPhone and iPad app that carries no account, no advertising and no in-app purchases.

Keep reading