Context windows: what the number really buys you
A million-token context window does not mean a model reads a million tokens well, or that you should send them. Here is what the number measures, what it costs, and when retrieval beats stuffing everything in.
Model announcements lead with the context window: 128,000 tokens, 200,000, a million. It is presented as a measure of how much the model can "remember", and the bigger number is implied to be the better model.
It is a real limit and worth understanding. But it measures less than it seems to, and the cost of using it sits on your side.
What the number counts
The context window is the maximum number of tokens the model can take into account in a single request. Everything counts against it:
- the system prompt,
- every earlier message in the conversation, because chat APIs are stateless and the whole history is sent again each turn,
- any documents, code or tool results you include,
- and, on most APIs, the output the model generates.
Many providers also set a separate, much smaller cap on output tokens. A model with a million-token window may still stop writing after a few thousand tokens of answer.
What a token is, roughly
A token is a piece of text the model's tokenizer treats as one unit — often a word, sometimes part of one. For ordinary English, a common rule of thumb is about three-quarters of a word per token, so 100,000 tokens is in the region of 75,000 words.
That rule is English-specific. Languages with rich morphology, such as Turkish or Finnish, and languages written in non-Latin scripts often need noticeably more tokens for the same meaning. Code, JSON and long numbers are also token-hungry. If you are budgeting, measure with your own content and the provider's tokenizer or token-counting endpoint rather than trusting a conversion factor.
Fitting is not the same as reading well
A model that accepts a million tokens will produce an answer about all of them. Whether it used the right parts is a separate question.
Research has repeatedly found that models are better at using information near the start and end of a long input than in the middle — the effect was described in the 2023 paper Lost in the Middle and has been measured in various forms since. Newer models have improved considerably, and providers publish long-context benchmarks to show it. But the benchmarks tend to test finding a planted fact, which is easier than reasoning across many scattered ones.
The practical reading: a large window makes long inputs possible. It does not guarantee they are handled as well as short ones, and you should test with your own documents before relying on it.
What it costs
This is the part the headline number hides.
You pay for input tokens on every request. If you put a 150,000-token document in the context and ask ten questions in a conversation, you have sent it ten times. At typical input prices, that is the difference between a fraction of a cent and a real line item.
Latency grows with input. A long prompt takes longer before the first token of the answer appears.
Prompt caching changes the maths. Most major providers now offer some form of caching for a repeated prefix — the same system prompt or document at the start of each request — billed at a reduced rate after the first time. If your long context is stable, structure requests so it comes first and unchanged, and check the provider's caching rules and minimums.
Some providers also price differently above a certain context length. Read the pricing page for the exact tier you would be in, not just the headline rate.
Stuffing or retrieval
The alternative to sending everything is retrieval: split your material into pieces, find the few relevant to this question, and send only those.
Send everything when the material is modest, the questions need the whole thing at once (summarising a contract, reviewing a single codebase module), and caching makes repetition cheap.
Retrieve when the collection is large or growing, each question touches a small part of it, or you are paying per request at volume.
Plenty of real systems do both: retrieval to pick the relevant documents, then a generous context to include them whole rather than in fragments.
How to compare windows honestly
When choosing between models, three numbers matter more than the maximum:
- The size of your typical request, measured in tokens with your real content.
- The input price at that size, including any long-context tier and any caching discount you would actually get.
- Quality at that size, tested on fifty real requests, not taken from a benchmark.
A model with a smaller window that comfortably fits your requests is not worse for your purposes. It may be cheaper and faster.
Where AIonRadar helps
AIonRadar follows model and API releases with links back to the original announcements and pricing pages, so a context limit or price you read there can be checked at the source rather than trusted second-hand.
Its API cost calculator lets you put in your own token volumes and see what a workload costs across models, which is where context length stops being a spec-sheet number and becomes a monthly bill. The LLM API Selector and source-backed comparisons are a quicker starting point than opening nine pricing pages.
It is free on the web, and the iPhone and iPad app has no account, no advertising and no in-app purchases.