← Notes

What benchmarks don't tell you

Every model launch comes with a table of benchmark scores, and the table is rarely what decides whether a model works for you. Here is what the numbers measure, how they mislead, and how to build a test that answers your actual question.

A new model arrives with a table: a dozen benchmarks, a column for each competitor, and the new model's scores in bold. It is the first thing people share and the last thing that should decide what you build on.

Benchmarks are not useless. They are narrow, and knowing exactly how narrow is what makes them useful.

What a benchmark is

A benchmark is a fixed set of questions or tasks with a way of scoring answers. Some are multiple choice across academic subjects. Some are mathematics problems with checkable answers. Some are programming tasks graded by running tests. Some ask people to compare two models' answers and pick the better one.

Each measures one kind of behaviour under one set of conditions. The score is a summary of that — no more.

Why the numbers mislead

Contamination. Benchmark questions are published, and models are trained on enormous amounts of text from the internet. If the questions — or discussions of them — were in the training data, the model may have seen the answers. Newer benchmarks try to avoid this with private or recently created questions, but a high score on an old public benchmark says less than it appears to.

Saturation. When the top models all score in the high 90s, the remaining gap is often noise, ambiguous questions or errors in the answer key. A benchmark near its ceiling no longer separates models.

Conditions differ. A score depends on the prompt format, the number of examples given, whether the model was allowed to "think" at length, how many attempts it had, and the settings used. Two companies reporting the same benchmark may not have run the same test. Footnotes are where this lives.

Self-reporting. Launch tables are produced by the company launching the model. That does not make them false, but independent re-runs sometimes produce different numbers, and the choice of which benchmarks to show is itself a claim.

Preference is not correctness. Leaderboards based on people voting between two answers measure what people prefer to read. Longer, more confident, better-formatted answers tend to win, whether or not they are more accurate.

Nothing about cost or speed. Benchmarks usually ignore price per request, latency, rate limits and reliability. A model two points better that costs five times as much, or answers three times more slowly, may be the wrong choice.

What they cannot tell you at all

How the model handles your data. Your documents, your terminology, your formats, your languages. A model that scores well on general knowledge can still misread your invoices.

Consistency. A benchmark reports how often the model is right on average. It does not tell you how badly it fails when it is wrong, or whether it fails the same way every time.

Following your instructions. Whether it keeps to a JSON schema, respects a length limit or stays in the tone you asked for across thousands of requests.

Behaviour over a long conversation or a long document, unless that specific benchmark tests it.

Building a test that answers your question

You do not need a research lab. You need a small, honest evaluation set:

  1. Collect 30 to 100 real examples from your own use — actual inputs, not invented ones. Include the awkward cases.
  2. Write down what a good answer looks like for each. For extraction tasks, the exact expected values. For open-ended tasks, a short checklist.
  3. Run every candidate model on the same set, with the same prompt and settings.
  4. Score them — automatically where the answer is checkable, by hand or with a clear rubric where it is not.
  5. Record cost and latency alongside the quality score.
  6. Keep the set. When the next model arrives, rerunning it takes minutes, and it is a better guide than any launch table.

Look at the failures, not just the score. A model that fails rarely but badly can be worse than one that fails more often in harmless ways.

Reading the launch table sensibly

Benchmarks are still useful for a first cut. Use them to decide which three models to test, not which one to choose. Prefer benchmarks that resemble your task, that were run independently, and that were published recently enough to have avoided contamination. Read the footnotes.

Where AIonRadar helps

AIonRadar follows model releases with links back to the original announcements, technical reports and model cards — the places where the conditions behind a score are actually stated. Its comparisons are source-backed rather than copied from launch tables.

The LLM API Selector narrows the field to models that fit a described workload, and the API cost calculator puts a monthly price next to each candidate, which is the column the benchmark table leaves out.

It is free on the web, and the iPhone and iPad app has no account, no advertising and no in-app purchases.

Keep reading