LLM Handbook

#Glossary

Definitions are one line. The second line is the part that matters — the consequence, the trap, or the reason the term exists. Terms are grouped by where they belong rather than alphabetically, because that is how they are used.

Use the search box (press /) to jump to a term.


#Architecture & model internals

TermDefinitionWhat matters about it
TransformerNeural architecture built on self-attentionTwo costs in different places: attention is quadratic in context, the MLP holds most of the weights
Self-attentionEach token attends to every other, weighted by relevanceThe softmax(QKᵀ/√d)V step; the √d stops the softmax saturating
Query / Key / ValueThe three projections attention computes from each tokenQuery is what a token seeks, key what it offers, value what it contributes
Multi-head attentionSeveral attention operations in parallel with different projectionsHeads specialise; many contribute almost nothing, which is what makes pruning possible
KV cacheStored keys and values for previous tokensGrows linearly with context and batch; at long context it exceeds the model weights
GQA (grouped-query attention)Several query heads share one key/value headShrinks the KV cache several-fold with little quality loss; now standard
MQA (multi-query attention)All query heads share one KV headGQA's more aggressive predecessor; more saving, more quality cost
FlashAttentionIO-aware attention implementationChanges memory access, not mathematics — identical output, much faster
RoPERotary position embeddingEncodes relative position; why context windows can be extended after training
ALiBiLinear attention bias by distanceAlternative position scheme that extrapolates cheaply
MoE (mixture of experts)Route each token to a few expert MLPsTotal parameters grow while per-token compute stays flat
Decoder-onlyCausal attention; each token sees only what precedes itWon because training parallelises over every position, not for representational reasons
Encoder-onlyBidirectional attentionEmbeddings, classification, rerankers — BERT-family
LogitsRaw pre-softmax scores over the vocabularyWhat the model actually outputs; decoding turns them into a token
PrefillProcessing the whole prompt at onceCompute-bound; sets time-to-first-token
DecodeGenerating one token at a timeMemory-bandwidth-bound; why quantization speeds it up

Interactive simulation — needs JavaScript.


#Tokenization

TermDefinitionWhat matters about it
TokenA subword unit; the model's atomic inputThe model never sees characters, which is why it cannot count letters
BPEByte-pair encoding; merges frequent adjacent pairsFrequency-derived, so English fragments least and other scripts most
VocabularyThe fixed set of tokens, typically 32k–200kCoupled to the model; counting with the wrong tokenizer gives wrong estimates
Special tokensBOS, EOS, chat role markersConsume budget and are easy to forget when counting
Glitch tokenA vocabulary entry barely present in trainingIts embedding is essentially arbitrary; behaviour on it is undefined, not just poor
Context windowMaximum tokens the model acceptsAdvertised length exceeds usable length; recall degrades before the limit, worst in the middle

#Decoding

TermDefinitionWhat matters about it
TemperatureDivides logits before softmax0 for extraction and judging, ~0.7 for chat. 0 is not determinism
Top-kKeep the k highest-probability tokensFixed regardless of confidence — wrong in both directions
Top-p / nucleusKeep the smallest set reaching probability mass pAdapts to confidence; the one to use
Greedy decodingAlways take the highest-probability tokenRepetitive; the temperature-0 case
Beam searchKeep b best partial sequencesWins for translation, loses for chat — likely text is bland text
Frequency / presence penaltyDown-weight tokens already usedBlunt: harmful on code, where repetition is correct
Speculative decodingDraft model proposes, large model verifies2–3× faster with mathematically identical output, so no re-evaluation needed
Constrained decodingA grammar masks invalid tokensMakes schema violation impossible, not merely unlikely
LogprobsPer-token log-probabilitiesA free, uncalibrated confidence signal — good for routing, not for showing users

#Retrieval & RAG

TermDefinitionWhat matters about it
RAGRetrieve documents, put them in context, generateFixes knowledge and freshness; does not fix behaviour
ChunkA retrievable piece of a documentRetrieval can only return a chunk, so a chunk without the answer is a hard ceiling
OverlapRepeating tokens across chunk boundariesBuys boundary safety at permanent storage cost of C/(C−O)
Small-to-bigRetrieve on small chunks, send the parent sectionSmall embeds precisely, large answers well; decoupling gets both
EmbeddingVector representation where similar text is nearbyCosine scores have no absolute meaning — rank with them, do not threshold blindly
Bi-encoderEncodes query and document separatelyPrecomputable, scales to millions, cannot model interaction
Cross-encoderEncodes query and document togetherFar more accurate, one model pass per pair, unusable over a corpus
Late interaction / ColBERTVector per token, scored by MaxSimBetween the two: precomputable docs, token-level precision, large storage
ANNApproximate nearest-neighbour searchApproximate — measure recall against an exact index or you have an unknown ceiling
HNSWLayered navigable small-world graph indexef_search is the live recall/latency dial; M costs memory
IVF / PQCluster-based index / compressed vectorsFor very large sets; PQ trades noticeable recall for memory
Hybrid retrievalBM25 plus dense, fusedEach fails where the other succeeds
BM25Classic lexical ranking functionExact terms, names, codes — where embeddings are weakest
RRFReciprocal rank fusion, 1/(k+rank) summedFuses by rank because scores from different systems are not comparable
RerankingReordering a shortlist with a stronger modelOptimises precision; retrieval optimises recall. Different objectives
HyDEEmbed a hypothetical answer instead of the questionWorks because the fake answer may be wrong — it is a probe, never shown
Query transformationRewriting the query before retrievingFixes question/answer asymmetry; route it, do not apply it to everything
Corrective RAGGrade retrieved passages, act if none are relevantAdds the branch plain RAG lacks: the ability to decline
Self-RAGAlso grades its own outputCatches drift beyond good context, a different failure from bad retrieval
Retrieval ceilingFraction of answers present in any chunkThe hard upper bound on the whole pipeline; measure it first

#Evaluation

TermDefinitionWhat matters about it
Golden setFrozen, versioned inputs with referencesFreeze it, or today's score is not comparable to yesterday's
BaselineThe last accepted scores, committedA score with nothing to compare against is a vanity number
GateThe rule that fails the buildUngated evals are ignored within about three weeks
hit@kAny relevant chunk in the top kCoarse; says nothing about ranking
recall@kFraction of relevant chunks retrievedDiffers from hit@k only for multi-chunk answers
MRRMean reciprocal rank of the first relevant resultRank-sensitive; catches improvements hit@k cannot see
nDCGDiscounted cumulative gain, normalisedFor graded relevance rather than binary
GroundednessWhether the answer is supported by the contextFor retrieval-only systems, asked one step earlier: is the answer there at all
LLM-as-judgeUsing a model to score outputsUnvalidated, it is a number with a decimal point — measure agreement with humans
Cohen's kappaAgreement corrected for chanceBelow 0.4 the judge is noise; raw agreement flatters imbalanced sets
Position biasJudges prefer whichever answer came firstSwap the order and require both to agree
Verbosity biasJudges reward longer answersReport length alongside score
Wilson intervalConfidence interval for a proportionAt n=50, ±8–11 points — which is why a 4-point move is not evidence
McNemar / paired comparisonCompare per-item outcomes on the same itemsCancels shared variance; the right way to compare two runs
DriftThe world moves while the model does notFour kinds — inputs, rule, corpus, vendor — with four different remedies
Covariate driftP(X) changesInputs look different; the model may still be right
Concept driftP(Y|X) changesSame input, different correct answer — the model is now wrong
PSIPopulation stability indexConventional bands (0.1 / 0.25) are folklore with a useful shape
DisaggregationComputing metrics per segmentAn aggregate is an average over people and hides who it fails

#Optimization & serving

TermDefinitionWhat matters about it
QuantizationStoring weights in fewer bitsSpeeds up decode because it is bandwidth-bound; the arithmetic is unchanged
PTQ / QATPost-training / quantization-aware trainingPTQ takes minutes and is right for almost everyone
GPTQ / AWQSecond-order and activation-aware 4-bit methodsBoth are responses to emergent outlier features
Outlier featuresActivation dimensions 10–100× larger, above ~6.7BWhy naive INT8 activation quantization destroys large models
DistillationTraining a small student to imitate a large teacherSoft labels carry "dark knowledge"; most teams actually do synthetic data generation
PruningRemoving weightsUnstructured gives no speedup without sparse kernels
Structured / 2:4 sparsityRemoving whole units / 2 of every 4 weightsThe forms that actually make inference faster
Continuous batchingScheduling at each token stepA finished sequence leaves immediately; several times static batching's throughput
PagedAttentionKV cache paged like virtual memoryCuts waste, enabling much larger batches in the same VRAM
TTFTTime to first tokenWhat users judge; set by prompt length and queueing
TPOTTime per output tokenSet by bandwidth and batch size; a different problem from TTFT
Prefix cachingReusing computation for a shared prompt prefixOften the largest single cost lever on a stable system prompt
Load sheddingRejecting work above a thresholdNeeds hysteresis — two thresholds — or it flaps
BackpressureSignalling upstream to slow downBounded queue plus 429 with Retry-After
Dynamic batching (alias)Forming a batch from whatever has arrived, rather than a fixed sizeOften used loosely for continuous batching — worth asking which is meant, because the per-request and per-step versions have very different tail latency
In-flight batching (alias)NVIDIA's name for continuous batchingSame mechanism, TensorRT-LLM's vocabulary. See Continuous batching
Paged KV cache (alias)The cache organised into fixed-size blocks with a block tableThe storage half of PagedAttention; the kernel half is what reads it
KV cache evictionDropping low-value entries to hold a fixed budgetIrreversible and input-dependent — an evicted token cannot be recovered if a later query needed it
KV cache token pruning (alias)Eviction chosen by a saliency scoreSame operation, named for how the victim is picked
KV cache sparsity (alias)Keeping or reading only part of the cacheTwo different things wear this name: sparse storage lowers memory, sparse reads lower bandwidth and do not
Kernel tilingProcessing in cache-sized blocks instead of whole rowsThe transformation underneath FlashAttention — it is why the N×N score matrix never exists
Fused prologueFolding work into a matmul before it, rather than afterThe mirror of a fused epilogue; less common, because most elementwise work lands downstream
Mixed-precision quantizationDifferent bit widths for different parts of one modelOutlier channels or sensitive layers stay wide while the rest go narrow — the practical middle ground between uniform INT4 and giving up
Blockwise quantizationOne scale factor per block of 32 or 128 weightsWhere almost everything has landed: fine enough to contain an outlier, coarse enough that the scales stay small
Cross-attentionAttention where queries come from one sequence and keys/values from anotherThe encoder–decoder link, and how a vision encoder's output reaches a language model
Dynamic inference (adaptive inference)Spending different amounts of compute per inputThe umbrella over early exit, MoE routing and cascades — the unifying idea is that not every input deserves the same work
Inference budgetA per-request cap on tokens, time or costWhat turns adaptive inference from a research idea into an SLO; without one, "spend more on hard inputs" has no ceiling
Deep learning compilerAhead-of-time optimiser for a model graph — TVM, XLA, InductorMost of what it does is operator reordering and fusion; it is why hand-written kernels are worth it only where it fails
Neural architecture search (NAS)Searching automatically over model shapesLargely superseded for LLMs by scaling-law extrapolation, which predicts the answer more cheaply than search finds it
Chunked attentionAttention computed over fixed blocks of the sequenceThe tiling idea applied to the sequence axis; the basis of block-sparse and sliding-window kernels
QKV computationProducing the query, key and value projectionsPacked into one GEMM in practice — three matrices with the same input and shape is one matmul, not three
Mixture-of-Heads (MoH)MoE routing applied to attention heads instead of FFN expertsPer-token head selection; the same conditional-computation idea moved from width to attention
Length generalizationWhether a model works past the length it trained onThe property RoPE scaling is trying to buy, and the reason a context window is a claim about memory rather than recall
Long RAGRetrieving fewer, larger passages for a long-context modelThe middle position between many small chunks and stuffing the corpus — the trade moves with the model's real recall, not its window
RAG cacheCaching retrieval results, or the KV of retrieved chunksTwo different things share the name: caching the documents saves the retriever, caching their KV saves prefill
RAG fusionRunning several query variants and fusing the ranked listsUsually reciprocal rank fusion; buys recall for the cost of N retrievals
Speculative RAGDrafting an answer before, or in parallel with, retrievalSame bet as speculative decoding — cheap when the draft is usually right, wasted work when it is not
Fake quantization (simulated quantization)Rounding to a quantized grid but computing in floatWhat QAT does during training, and what a benchmark is doing when INT4 shows no speedup — it measured accuracy, not throughput
Stochastic quantizationRounding up or down with probability set by the remainderUnbiased in expectation, which matters for gradients and rarely for inference
Cluster-based quantization (weight clustering)Weights replaced by indices into a learned codebookVector quantization's ancestor; AQLM is the modern form
Block floating-pointA shared exponent across a block, with per-element mantissasThe format underneath MXFP4 and NVFP4 — "block-scaled" describes the scaling, "block floating-point" the representation
FTZ / DAZCPU flags that flush denormals to zeroDenormals can cost hundreds of cycles; these trade a sliver of precision near zero for predictable latency
Edge inferenceRunning on the device rather than a serverMemory and power bound, not FLOP bound — which is why edge models are quantized aggressively and often distilled
AI PCA laptop or desktop with an on-board NPUThe consumer form of edge inference; the NPU is built for low-precision integer throughput at low power
Serverless inferencePer-request compute with no persistent instanceCold starts dominate: loading weights is the cost, so it suits spiky low-volume traffic and suits nothing that needs a warm KV cache

#Agents, adaptation & safety

TermDefinitionWhat matters about it
AgentA loop where the model chooses the next actionEvery hard problem here is loop control
ReActInterleaved reasoning and actingThe reasoning step is what enables recovery from surprises
Tool / function callingModel emits a structured call against a schemaTool descriptions and error messages decide quality more than model choice
TrajectoryThe sequence of steps takenScore it separately from the outcome; a lucky path will not generalise
MCPModel Context ProtocolA standard way to expose tools and resources to models
LoRALow-rank adapters over frozen weights~0.1% of parameters; adapters swap per request
QLoRALoRA over a 4-bit quantized baseFine-tuning on consumer hardware
Catastrophic forgettingLosing general ability while gaining a narrow oneRoutinely unmeasured; evaluate a general set before and after
RLHF / DPOAligning to human preferencesRLHF uses a reward model; DPO optimises preferences directly and is simpler
Prompt injectionInstructions smuggled in as dataThe model cannot separate instructions from data; controls must be architectural
Indirect injectionInjection via retrieved contentNeeds no malicious user; with tools it becomes exfiltration
Dual-LLM patternPrivileged model never reads untrusted textThe strongest structural mitigation; restrictive but addresses the cause
Least privilegeTools can do only what is neededBounds the blast radius when injection succeeds — and assume it will
Denial of walletDriving cost rather than stealing dataRate limits and max_tokens are security controls, not just capacity ones

#Business

TermDefinitionWhat matters about it
Contribution marginRevenue minus variable cost per unitLLM cost scales with usage, so heavy users can be loss-making
Build vs buyHosted API versus self-hostingDecided by hidden costs — on-call, upgrades, expertise — not per-token price
MoatWhat competitors cannot copyProprietary data, workflow lock-in, distribution, regulation. Not prompts or model choice
Batch APIAsynchronous processing at reduced priceRoughly half price where latency permits
Shadow modeRun the new system without showing outputThe correct first rollout step; real traffic, zero user risk
CanarySmall percentage of live trafficCatches what offline evaluation cannot

#How to use this page

Every entry here is expanded somewhere in the handbook. If a one-line definition is not enough, the page that covers it properly is one search away — and if a term appears here that you cannot yet explain to a sceptical interviewer, that page is where to go next.