Quantisation, the practical version
Quantisation is why a model that "needs" 140 GB runs on a laptop. Here is what it does, what the file names like Q4_K_M mean, how much quality you give up at each level, and how to choose.
Download an open-weights model to run locally and you are immediately faced with a list of files: Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q3_K_S, IQ2_XS, or 4bit, 8bit, AWQ, GPTQ. They are the same model at different levels of quantisation, and the one you pick decides whether it fits in memory, how fast it runs and how well it answers.
We have covered how much VRAM a model needs before. This is the practical companion: what the options mean and how to choose between them.
What quantisation does
A model's weights are numbers, and they are usually trained and published as 16-bit values — two bytes each. A 70-billion-parameter model at 16 bits is about 140 GB.
Quantisation stores each weight with fewer bits:
- 8 bits — about half the size: ~70 GB for that model.
- 4 bits — about a quarter: ~35–40 GB.
- 2–3 bits — smaller still, with growing costs.
Fewer bits means less precision. The trick is that most weights do not need all sixteen bits to do their job, so a well-designed scheme loses surprisingly little.
Why smaller is also faster
Generating text on consumer hardware is mostly limited by memory bandwidth — how quickly the weights can be read for each token — rather than by raw computation. A model a quarter of the size can be read roughly four times as fast. So quantisation usually improves speed as well as fit, particularly on Apple silicon and consumer GPUs.
How much quality you lose
The honest answer is "it depends on the model and the task", but broad patterns are well established:
- 8-bit is close to indistinguishable from the original for most uses.
- 5- and 6-bit are very close, with small measurable differences.
- 4-bit, using a good modern scheme, is where most people land. Quality loss is noticeable in careful tests but modest in everyday use. This is usually the best trade-off.
- 3-bit shows clearer degradation — more errors, weaker reasoning, more repetition.
- 2-bit is a last resort. It can work for very large models; smaller models degrade quickly.
Two further patterns:
Larger models tolerate quantisation better. A 70B model at 4 bits usually loses less, relatively, than a 7B model at 4 bits.
A bigger model quantised often beats a smaller model unquantised. If you have a fixed amount of memory, a larger model at 4 bits is frequently better than a smaller one at 8 or 16 bits. Test both on your own task.
Some tasks are more sensitive. Mathematics, precise code, and following long, detailed instructions suffer earlier than general chat and summarisation.
Reading the file names
Different tools use different formats.
GGUF — used by llama.cpp and tools built on it, such as LM Studio and Ollama. Names like:
Q8_0— 8-bit, simple scheme.Q6_K,Q5_K_M,Q4_K_M,Q3_K_S— the number is roughly the bits per weight;Kmeans the newer "k-quant" method that gives important parts of the model more precision;S,MandLare small, medium and large variants of that mix.IQ4_XS,IQ3_M,IQ2_XS— "i-quants", a further refinement that keeps quality better at very low bit counts, sometimes at a speed cost on some hardware.
Q4_K_M is the common default recommendation, and a sensible starting point.
MLX — Apple's framework for Apple silicon. Models are published as 4bit, 6bit or 8bit variants, and run efficiently using the Mac's unified memory.
GPTQ and AWQ — schemes aimed mainly at NVIDIA GPUs, commonly used by inference servers.
FP8 and similar — 8-bit floating-point formats supported natively by recent data-centre GPUs, used for serving at scale.
The part people forget: context memory
Quantising the weights does not shrink the memory used for the context — the KV cache, which grows with the length of the conversation or document. At long context lengths it can take several gigabytes on its own.
Some tools can quantise the KV cache as well, at a small quality cost. If a model fits at short context and fails at long context, this is usually why.
How to choose
- Work out your memory budget. GPU memory, or on a Mac, roughly two-thirds to three-quarters of unified memory for a comfortable margin.
- Leave room for context — a few gigabytes at least, more for long documents.
- Start at 4-bit (
Q4_K_M, MLX4bit, or AWQ) of the largest model that fits. - Test on your own examples. Ten or twenty real prompts.
- Adjust. If quality is lacking, try a higher bit level or a larger model; if speed is lacking, go smaller.
Download from sources you trust. Quantised versions are often made by third parties rather than the original publisher; reputable ones state which base model and method they used, and the base model's licence still applies.
Where AIonRadar helps
AIonRadar's GPU and VRAM calculator estimates memory requirements for a model at different quantisation levels and context lengths, which is the first step in the routine above.
For the other side of the decision, the API cost calculator and LLM API Selector show what the same workload would cost through a hosted API — sometimes the more sensible answer. Its news follows model releases with links to original sources, including model cards and licences.
It is free on the web, and the iPhone and iPad app has no account, no advertising and no in-app purchases.