← Notes

Rate limits, and what they do to your product

The model worked perfectly in testing, then launch day brought a wall of 429 errors. Here is what LLM API rate limits measure, why they bite at the worst time, and how to design around them before users notice.

In development, the model answers every request. At launch, traffic arrives all at once and the API starts returning 429 Too Many Requests. Users see spinners, errors or nothing at all.

Rate limits are the least glamorous part of choosing an AI API and one of the most important. They decide how many people your product can serve, and they are easy to miss because they rarely appear on the pricing page.

What gets limited

Providers typically limit several things at once, and hitting any one of them stops you:

  • Requests per minute (RPM) — how many calls, regardless of size.
  • Tokens per minute (TPM) — how much text, often split into input and output tokens.
  • Requests or tokens per day — a daily ceiling on some tiers.
  • Concurrent requests — how many can be in progress at the same moment.

A small number of large requests can hit the token limit while barely touching the request limit. Many tiny requests can do the opposite.

Tiers: your limit depends on your history

Most providers assign limits by usage tier, and you move up tiers by spending more over time or by requesting an increase. A new account starts low.

Two consequences:

  • The limits you tested against may not be the limits you launch with. A personal test account and a production account can be on different tiers.
  • New models often have lower limits at launch. A just-released model may be limited more tightly than an established one, for everyone, until capacity grows.

Check the actual limits for your account and the exact model in the provider's console, not in a blog post.

How limits hurt a product

Spikes, not averages. An app that averages ten requests a minute may send two hundred in the minute after a newsletter goes out. Limits are enforced per minute or less, so the spike is what counts.

Long requests hold capacity. A request with a large context or a long answer uses a lot of the token budget, and occupies a concurrency slot for longer.

Retries make it worse. When everything fails and every client retries immediately, the retries become a second spike on top of the first.

Agents multiply calls. One user action in an agent-style feature can trigger many model calls — tool use, planning, checking. The user sees one button; the API sees ten requests.

Designing around them

Read the headers. Many APIs return headers saying how much of each limit remains and when it resets. Log them. They tell you how close you are before you hit the wall.

Retry with exponential backoff and jitter. On a 429, wait, then retry, doubling the wait each time, plus a small random amount so clients do not retry in lockstep. If the response includes a retry-after value, respect it.

Queue instead of failing. For anything that does not need an instant answer, put requests in a queue and process them at a rate below your limit. Users see "processing" rather than an error.

Use batch APIs for bulk work. Several providers offer batch endpoints for large jobs that can wait hours — typically at a discount and with separate, higher limits. Overnight summarisation or classification belongs there.

Cache what repeats. Identical requests should not hit the API twice. Prompt caching also reduces the cost of repeated prefixes, and on some providers it affects how requests count against limits.

Trim tokens. Shorter prompts, tighter context and sensible output length caps help both cost and token limits.

Have a fallback. A second model — a smaller one from the same provider, or one from another provider — that you can switch to when the first is limited or unavailable. Test it in advance; fallbacks that have never run tend not to work.

Limit your own users. Your per-user limits should be designed so that no single user, or script, can exhaust your API allowance for everyone else.

Estimating before launch

A rough calculation catches most surprises:

  1. Expected peak users per minute × model calls per user action = peak RPM.
  2. Peak RPM × average tokens per call = peak TPM.
  3. Compare both with your tier's limits for the exact model, with a safety margin of at least two.

If the numbers do not fit, request a limit increase well before launch — approval is not instant — or redesign so the peak is lower.

Where AIonRadar helps

AIonRadar follows AI model and API announcements with links to the original sources, which is where changes to tiers, model availability and batch or caching options are published.

Its API cost calculator turns your expected token volumes into monthly costs across models — the same numbers you need for the peak calculation above — and the LLM API Selector and source-backed comparisons help pick a primary model and a realistic fallback.

It is free on the web, and the iPhone and iPad app has no account, no advertising and no in-app purchases.

Keep reading