LLM Handbook

#Quantization

The one sentence: token generation is limited by how many bytes you can pull out of memory, not by how fast you can multiply — so making the weights smaller makes the model faster.

That sentence is the whole topic. Almost everyone can define quantization; far fewer can say why it produces a speedup, and the answer is not "smaller numbers multiply faster".


#1 · Diagram

   WHY GENERATION IS SLOW: THE TWO PHASES ARE NOT ALIKE

   PREFILL (reading the prompt)          DECODE (writing each token)
   ----------------------------          ---------------------------
   all prompt tokens at once             one token at a time
   big matrix x big matrix               big matrix x ONE vector
   arithmetic dominates                  MOVING THE WEIGHTS dominates
   -> COMPUTE-BOUND                      -> MEMORY-BANDWIDTH-BOUND


   DECODE, PER TOKEN:  read every weight once, do very little maths with it

     7B model, FP16   =  14 GB of weights read PER TOKEN
     at 2 TB/s HBM    =  ~7 ms/token   ->   ~140 tokens/s ceiling

     same model, INT4 =  3.5 GB read per token
     at 2 TB/s        =  ~1.75 ms/token ->  ~570 tokens/s ceiling

   The maths did not get faster. There is simply 4x less to carry.

This is the reason quantization is the highest-leverage inference optimization, and the reason it helps decode far more than prefill.


Interactive simulation — needs JavaScript.


#2 · Design

Basic. Quantization maps high-precision floats to a smaller numeric type. A tensor of FP16 values becomes INT8 or INT4 plus a scale (and often a zero-point) so the original range can be approximately recovered.

   real value  ≈  scale × (quantized_int − zero_point)

The memory arithmetic everyone should be able to do in their head:

PrecisionBytes/param7B model70B model
FP32428 GB280 GB
FP16 / BF16214 GB140 GB
INT8 / FP817 GB70 GB
INT40.53.5 GB35 GB

Add roughly 15–20% for the KV cache, activations and framework overhead. That table is what decides whether a model fits on the card you have — a 70B at INT4 fits on a single 40GB A100; at FP16 it does not fit on two.

Intermediate — the axes that actually distinguish methods.

AxisOptionsConsequence
What is quantizedWeights only / weights + activationsWeight-only is far easier and safer; activations carry outliers
WhenPost-training (PTQ) / during training (QAT)PTQ takes minutes; QAT needs a training run
GranularityPer-tensor / per-channel / per-groupFiner granularity, better accuracy, slightly more overhead
SymmetrySymmetric / asymmetricAsymmetric handles skewed distributions, costs a zero-point

Advanced — the finding that shaped the whole field. Above roughly 6.7B parameters, transformers develop emergent outlier features: a small number of dimensions whose activations are 10–100× larger than the rest. Naive INT8 quantization of activations destroys the model, because those outliers dominate the scale and everything else collapses into a couple of quantization levels.

The responses to that discovery are the modern method landscape:

MethodCore ideaWhere it fits
LLM.int8()Keep outlier dimensions in FP16, quantize the rest to INT8The original fix; simple, some overhead
GPTQLayer-by-layer, second-order-informed rounding against calibration dataStrong 4-bit accuracy, needs calibration
AWQProtect the ~1% of weights that matter most, guided by activation scaleFast, robust, popular for 4-bit serving
SmoothQuantShift difficulty from activations into weights so both quantize wellEnables weight+activation INT8
GGUF / llama.cpp k-quantsMixed per-block precision, CPU-friendlyLocal and CPU inference
FP8Hardware-native float8 on H100 and newerNear-lossless, needs the silicon

Step 4 above says "run the quantizer" in one line. Here it is, actually running: fit a scale and zero-point to a vector you can edit, round every weight to an integer code, and read what the round-trip cost you.

Interactive experiment — needs JavaScript.

#3 · Flow

Post-training quantization, in the order it happens:

  1. Pick the target. Driven by the memory arithmetic above and the hardware you actually have, not by what sounds impressive.
  2. Choose weight-only or weight+activation. Weight-only INT4 is the default for serving; it captures most of the bandwidth win with much less risk.
  3. Gather calibration data — typically 128–512 samples. This is the step people rush, and it decides the result.
  4. Run the quantizer. GPTQ or AWQ for 4-bit; minutes to a couple of hours.
  5. Evaluate on your task, not on perplexity alone. See below.
  6. Compare against the alternatives you did not take — a smaller model at FP16 is often better than a large one at INT4.
  7. Deploy and measure real throughput, because the theoretical speedup assumes you were bandwidth-bound in the first place.
graph TD
  A[FP16 checkpoint] --> B{Fits in memory<br/>at acceptable speed?}
  B -->|yes| C[Ship it. Do nothing]
  B -->|no| D[Pick target precision from the memory table]
  D --> E[Gather calibration data<br/>from YOUR distribution]
  E --> F[Quantize: GPTQ / AWQ / FP8]
  F --> G[Evaluate on task metrics<br/>not perplexity alone]
  G --> H{Quality acceptable?}
  H -->|no| I[Coarser bits, finer granularity,<br/>or keep sensitive layers in FP16]
  I --> F
  H -->|yes| J[Benchmark real throughput]
  J --> K{Actually faster?}
  K -->|no| L[You were not bandwidth-bound.<br/>Look at batching instead]
  K -->|yes| M[Deploy]

Branch K → L catches a common disappointment. With large batches, decode stops being purely bandwidth-bound because each weight read is amortised across many sequences. Quantizing a heavily-batched server can produce a much smaller speedup than the memory arithmetic promised — the win is largest at batch size 1, which is exactly the interactive, single-user case.


#4 · UML — where quantization sits

graph LR
  subgraph Offline["Offline, once"]
    A[FP16 checkpoint] --> B[Calibration set]
    B --> C[Quantizer<br/>GPTQ / AWQ]
    C --> D[Quantized weights<br/>+ scales]
  end
  subgraph Serving["Per request"]
    D --> E[Load into VRAM]
    E --> F[Prefill: compute-bound]
    F --> G[Decode loop:<br/>bandwidth-bound]
    G --> H[Dequantize block<br/>to FP16 in-kernel]
    H --> I[Matmul]
    I --> G
  end

Note H. In weight-only quantization the arithmetic still happens in FP16 — weights are dequantized inside the kernel, block by block, as they are read. Nothing about the multiply got cheaper. The saving is entirely in the bytes crossing the memory bus, which is why the speedup tracks the compression ratio rather than any change in FLOPs.


#5 · Example

def memory_footprint(params_b: float, bits: int, kv_overhead: float = 0.18) -> float:
    """Approximate VRAM in GB for a model at a given precision.

    The overhead term covers the KV cache, activations and framework slack.
    It grows with context length and batch size, so treat 0.18 as a floor for
    interactive use rather than a promise.
    """
    weights_gb = params_b * 1e9 * (bits / 8) / 1e9
    return weights_gb * (1 + kv_overhead)


def decode_ceiling(params_b: float, bits: int, bandwidth_tb_s: float = 2.0) -> float:
    """Upper bound on tokens/second at batch size 1, from bandwidth alone.

    Every weight is read once per token, so the ceiling is simply
    bandwidth / bytes-per-pass. Real throughput lands below this -- attention
    over the KV cache, kernel launches and sampling all cost time -- but if
    your measured rate is far below this number, quantization is not your
    bottleneck and you should look elsewhere.
    """
    bytes_per_pass = params_b * 1e9 * (bits / 8)
    return (bandwidth_tb_s * 1e12) / bytes_per_pass


for bits in (16, 8, 4):
    print(f"7B @ INT{bits:<2}  {memory_footprint(7, bits):5.1f} GB   "
          f"ceiling ~{decode_ceiling(7, bits):5.0f} tok/s")
7B @ INT16   16.5 GB   ceiling ~  143 tok/s
7B @ INT8     8.3 GB   ceiling ~  286 tok/s
7B @ INT4     4.1 GB   ceiling ~  571 tok/s

Calibration data is the part that decides quality:

def calibration_samples(production_logs, n: int = 256) -> list[str]:
    """Draw calibration data from real traffic, stratified by intent.

    The quantizer decides which weights matter by watching activations on this
    data. Calibrate on WikiText and deploy on customer support transcripts and
    you have optimised for the wrong distribution -- perplexity will look fine
    and your task metrics will not.

    A few hundred samples is enough; representativeness beats volume.
    """
    by_intent: dict[str, list[str]] = {}
    for record in production_logs:
        by_intent.setdefault(record["intent"], []).append(record["text"])

    per_bucket = max(1, n // max(1, len(by_intent)))
    out: list[str] = []
    for texts in by_intent.values():
        out.extend(texts[:per_bucket])
    return out[:n]

#6 · Depth — the senior layer

Outlier handling and number formats are two answers to one problem, and the formats are winning. Everything in the outlier table above — LLM.int8(), AWQ, SmoothQuant — exists because a single scale factor for a whole tensor cannot express a distribution with extreme values in it. Block-scaled formats attack that in the number system instead: give every small block its own scale, and an outlier costs you its block rather than the tensor. That is why 4-bit became practical when NVFP4 and MXFP4 arrived rather than when the algorithms did. The two still compose — a rotation before a block-scaled quantizer still helps — but the algorithmic work is no longer carrying the whole burden. See kernel & attention optimization for the formats themselves, and for why a quantization scheme without a fused kernel is a paper rather than a deployment.

"Lossless" is used for two different things, and only one of them is. Worth separating before the word appears in a vendor deck.

Bit-exact lossless is compression, not quantization. DFloat11 is the clearest example: BF16 weights have low entropy in their exponent bits, so Huffman-coding those bits shrinks a model by about 30% with outputs bit-for-bit identical to the original. Weights stay compressed in VRAM and are decompressed on the fly before each matmul by a custom kernel using lookup tables in SRAM. There is no accuracy question to ask, because there is no accuracy change — the only costs are decode latency and kernel complexity. Against CPU offloading, the reported throughput gain is 2.3–46.2×, which is the comparison that matters: the real alternative to "model does not fit" is usually offload, not a smaller model.

Statistically lossless is a different and weaker claim: the output distribution is preserved within some tolerance, not reproduced exactly. That is often a perfectly good trade, but it is not the same promise, and it needs the same evaluation any lossy method needs.

The practical rule: if someone says lossless, ask whether they mean bit-identical or statistically indistinguishable. The first needs no eval. The second needs all of them.

Perplexity is a bad acceptance test and it is the one everybody uses. Quantization papers report perplexity because it is cheap and comparable, but a 0.1 perplexity increase can hide a large drop in a specific capability. The capabilities that degrade first, in roughly this order:

  1. Long-context recall — retrieval from deep in the context window
  2. Structured output — JSON validity, schema adherence, tool-call formatting
  3. Multi-step reasoning — errors compound across steps
  4. Rare languages and domain jargon — thin training signal, quantized away first

If your product depends on tool calls or strict JSON, test exactly that. A model that has quietly lost 5% JSON validity is a production incident, and perplexity will not show it.

Quantize the KV cache too, and know that it is riskier. At long context the KV cache can exceed the weights. INT8 KV is usually safe; INT4 KV degrades long-context recall noticeably. Quantize weights first, measure, then consider the cache separately — KV cache optimization covers that side, including why keys tolerate less precision than values.

The comparison people forget to run. Before quantizing a large model, check the smaller model at full precision:

OptionMemoryTypical outcome
13B @ INT4~7 GBOften worse than the alternative below
7B @ FP16~14 GBUsually stronger on structured output and reasoning
13B @ INT8~14 GBUsually the best of the three

There is no universal answer, but "big model, aggressive quantization" is not automatically better than "smaller model, gentle quantization", and the assumption that it is costs teams real quality.

Failure modeSymptomFix
Wrong calibration distributionBenchmarks fine, production worseCalibrate on real traffic, stratified
Perplexity-only acceptanceSubtle capability loss shipsTask metrics, especially JSON and tool calls
Quantizing an under-batched serverExpected 4×, got 1.2×You were not bandwidth-bound; batch first
INT4 KV cache at long contextRecall from early context collapsesKeep KV at INT8 or FP16
Uniform bits across layersAvoidable quality lossKeep the first and last layers higher precision
Ignoring hardware supportINT4 kernels slower than FP16Check the kernels exist for your GPU before choosing

Hardware determines what is even worth trying. FP8 needs Hopper or newer. INT4 needs good kernels — on hardware without them, dequantization overhead can make INT4 slower than FP16. The right order is: check what your silicon accelerates, then pick the format.

QAT versus PTQ, honestly. QAT gives better accuracy at very low bit widths (3-bit and below) but requires a training run, the data, and the expertise. For almost every application team, PTQ with GPTQ or AWQ at 4-bit is the right answer and QAT is a research project. Say that plainly rather than listing QAT as an equal option.


#7 · From each seat

SeatWhat quantization looks like from here
UserFaster responses, or the feature existing at all because it now fits on affordable hardware. The risk they carry is invisible: a subtly worse model that still sounds completely confident.
CoderPin the quantization method and its version alongside the model — a re-quantized checkpoint is a different model. Verify structured output after quantizing; that is what breaks first. Check kernel support before choosing a format.
TesterYour regression suite is the acceptance test for quantization. Run it on the quantized model and gate on task metrics, never perplexity. Add a JSON-validity and tool-call-accuracy check specifically — they degrade before anything a benchmark measures.
System designerChanges your memory budget, which changes batch size, which changes throughput — often more than the quantization itself. Model the whole chain, and measure at your real batch size, not at batch 1.
ArchitectIt is what makes self-hosting viable against an API. That is a build-versus-buy decision with a three-year shape: you take on serving, GPU capacity, and a quantization pipeline in exchange for per-token cost and data residency.
CEORoughly 4× less GPU memory means roughly 4× fewer GPUs for the same load, and the option to self-host at all. The risk is a quality loss nobody measured. Ask one question: what did the eval suite say before and after?
MarketCommoditised and moving fast — GPTQ, AWQ, GGUF, bitsandbytes are free and well-supported; vLLM and TensorRT-LLM handle serving. No moat. The moat is having an eval suite good enough to know whether the quantized model is still acceptable.

#8 · Interview questions

QuestionWhat they are testingAnswer sketch
"Why does quantization make inference faster?"The central misconceptionDecode is memory-bandwidth-bound: every weight is read once per token and very little arithmetic is done with it. Fewer bytes per pass means more tokens per second. The multiply is unchanged — weights are dequantized in-kernel.
"INT8 or INT4?"Trade-off reasoningStart at INT4 weight-only for serving; it captures most of the bandwidth win with modern methods. Go to INT8 if task metrics drop, especially structured output. And check whether a smaller model at higher precision beats both.
"What breaks first when you quantize?"Depth beyond the definitionLong-context recall, structured output validity, multi-step reasoning, and rare languages. Perplexity barely moves for all of these, which is why perplexity is the wrong acceptance test.
"Why do large models quantize worse than small ones?"Whether you know the outlier resultEmergent outlier features above roughly 6.7B: a few activation dimensions run 10–100× larger and dominate the scale. LLM.int8, SmoothQuant and AWQ are all responses to that.
"You quantized to INT4 and got a 1.2× speedup, not 4×."Systems thinkingYou were not bandwidth-bound. Large batches amortise weight reads across sequences, so the win shrinks. Check batch size and whether you are prefill-heavy — quantization helps decode far more than prefill.
"How much calibration data?"Practical knowledgeA few hundred samples, drawn from your real distribution and stratified by intent. Representativeness matters far more than volume; calibrating on generic web text and serving a domain workload is the classic mistake.

#Stop condition

You are done when you can:

  1. explain the speedup in terms of memory bandwidth, not arithmetic,
  2. do the memory arithmetic for a 7B and a 70B at four precisions from memory,
  3. name the outlier-feature result and two methods that respond to it,
  4. list what degrades before perplexity does, and
  5. state the case where a smaller model at FP16 beats a bigger one at INT4.

#Sources worth reading

TopicSource
Outlier featuresLLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (Dettmers et al., 2022)
4-bit PTQGPTQ (Frantar et al., 2022) and AWQ (Lin et al., 2023)
Activation quantizationSmoothQuant (Xiao et al., 2022)
ServingvLLM and TensorRT-LLM docs — which formats are actually accelerated on which hardware
Honest evaluationAny careful study reporting task metrics rather than perplexity alone; the gap between the two is the point