How much VRAM you need to run a model locally
The parameter count tells you roughly nothing on its own. What decides whether a model fits is the quantisation, the context length you want, and the overhead nobody includes in the arithmetic.
"Can I run a 70B model on a 24GB card?" is the most-asked question in local inference, and the answer is "it depends on three things nobody mentions when they quote the parameter count".
The arithmetic that gets you close
Model weights take up memory in proportion to the number of parameters and the number of bits used to store each one.
At 16-bit precision, each parameter needs 2 bytes. A 7-billion-parameter model is therefore about 14GB of weights. A 70B model is about 140GB. This is why nobody runs large models at full precision on consumer hardware.
Quantisation reduces the bits per parameter:
| Precision | Bytes per parameter | 7B | 13B | 70B |
|---|---|---|---|---|
| FP16 | 2 | ~14GB | ~26GB | ~140GB |
| 8-bit | 1 | ~7GB | ~13GB | ~70GB |
| 5-bit | ~0.65 | ~4.5GB | ~8.5GB | ~46GB |
| 4-bit | ~0.5 | ~3.5GB | ~6.5GB | ~35GB |
The rule of thumb: parameters in billions × bytes per parameter = weights in GB. Everything else is on top of that.
The three things that break the estimate
1. The KV cache, which grows with context. As the model generates, it stores keys and values for every token in the conversation. This is not fixed — it scales with context length and can become larger than you expect. A long-context session on a mid-sized model can add several gigabytes, and it grows as the conversation does. A model that loads fine and then runs out of memory at message forty is almost always this.
2. Overhead. The runtime, the CUDA or Metal context, activations during a forward pass, and the operating system's own use of the GPU. Budget roughly 1–2GB beyond the weights and cache, more on a machine that is also driving a display.
3. Quantisation is not free. 8-bit is generally indistinguishable from 16-bit for most tasks. 4-bit is usually usable and noticeably weaker on reasoning, code and long-context coherence. Below 4-bit, degradation becomes obvious. "It fits" and "it is worth running" are different questions.
So the practical formula is:
weights + KV cache for your target context + ~2GB overhead
A 13B model at 4-bit with a modest context fits comfortably in 12GB. A 70B model at 4-bit needs around 40GB once you account for everything, which means two 24GB cards or a large unified-memory Mac.
Apple silicon is a different calculation
On a Mac, the GPU uses unified memory shared with the system. There is no separate VRAM figure — a 64GB Mac can allocate a large share of that to a model, which is why Apple silicon machines punch well above their price for local inference of big models.
The trade is bandwidth. Memory bandwidth, not capacity, is what limits token generation speed, and a Mac that can hold a 70B model will generate noticeably more slowly than a dedicated GPU that can also hold it. Fitting and being pleasant to use are separate thresholds.
By default macOS reserves a portion of memory for the system; the usable share for the GPU is high but not the full amount.
When local is the wrong answer
Be honest about this before buying hardware:
- Low, bursty usage. If you make a few hundred calls a month, an API costs a few dollars and the GPU costs several hundred.
- You need frontier quality. The best open-weights models are very good. They are not the best models available.
- You need it to be fast on long prompts. Prompt processing on consumer hardware is slow compared to hosted inference.
Local wins on: data that must not leave your machine, high steady volume where per-token cost accumulates, offline requirements, and the ability to run something that will not be deprecated out from under you.
Where AIonRadar helps
AIonRadar has a GPU and VRAM calculator for exactly this arithmetic — you give it the model and the configuration you want and it works out the requirement, rather than you doing the table above by hand and forgetting the KV cache.
Alongside it, there is a calculator for API usage cost, which is the other half of the decision: the honest comparison is total cost of local hardware against the API bill for your real volume, not one in isolation.
It tracks model and hardware releases with links to original sources, so the numbers come from the people who published them. Free on the web, and the iPhone and iPad app has no account, no advertising and no in-app purchases.