LLM Handbook

#Interview question bank

Every question from across the handbook, consolidated. Each links back to the page that answers it properly.

⭐ marks the questions that most reliably separate candidates — usually because the obvious answer is wrong, or because most people stop one level short.


#How to drill this

Do not read the answers. Cover them, answer out loud, then check.

The failure mode of a question bank is recognition without recall: you read the answer, it feels familiar, you conclude you know it, and then you cannot produce it under pressure. Say it out loud — the gap between "I know this" and "I can say this in sixty seconds" is where interviews are lost.

Time-box to two minutes per question. Most answers should take sixty to ninety seconds. If you cannot finish in two minutes you are including material that is not earning its place.


Interactive simulation — needs JavaScript.


#Round 1 · Foundations

Asked to check you understand the machine, not just the API.

QuestionPage
Explain attention.Transformers
⭐ Why is decode memory-bound but prefill compute-bound?Transformers
What is the KV cache and why does it matter?Transformers
Why decoder-only for generation?Transformers
What does FlashAttention change?Transformers
⭐ Why can't the model count the letters in a word?Tokenization
Why is non-English text more expensive?Tokenization
Temperature 0 versus 0.7 — when each?Decoding
⭐ Is temperature 0 deterministic?Decoding
top_k or top_p, and why?Decoding
Why don't chat models use beam search?Decoding
What is speculative decoding?Decoding

#Round 2 · RAG and retrieval

The largest section, because it is the most-asked area.

QuestionPage
How would you chunk a technical manual?Chunking
⭐ RAG quality is poor. Where do you look first?Chunking
What is small-to-big retrieval?Chunking
⭐ How do you handle permissions in RAG?Chunking
You are changing embedding model. What is involved?Chunking
How does HNSW work?Embeddings & vector DBs
⭐ What recall does your ANN index actually get?Embeddings & vector DBs
Is 0.82 cosine similarity good?Embeddings & vector DBs
Which vector database, and why?Embeddings & vector DBs
⭐ Why two retrieval stages instead of one?Reranking
How deep should the shortlist be?Reranking
When does a reranker not help?Reranking
⭐ What is HyDE and why does it work?Query transformation
HyDE hallucinates. Isn't that a problem?Query transformation
Multi-query or reranking — which fixes what?Query transformation
How do you fuse results from several queries?Query transformation
How do you stop RAG answering from irrelevant context?Corrective RAG
CRAG versus self-RAG?Corrective RAG
⭐ Your relevance grader is sometimes wrong. Is that a problem?Corrective RAG

#Round 3 · Evaluation

Where most candidates are weakest, and where a good answer is most noticeable.

QuestionPage
⭐ How do you know your LLM judge is any good?LLM as a judge
Pointwise or pairwise judging?LLM as a judge
⭐ The score went up 4 points on 50 items. Ship it?LLM as a judge
The vendor updated the judge model. What now?LLM as a judge
When would you not use an LLM judge?LLM as a judge
How do you stop LLM quality regressing?Regression gates
⭐ Why gate per bucket as well as overall?Regression gates
Should a green run update the baseline?Regression gates
Someone edited the corpus. What should the gate do?Regression gates
How do you know your gate works?Regression gates
How would you detect drift in a live system?Drift detection
Covariate versus concept drift?Drift detection
⭐ You have no labels. How do you monitor quality?Drift detection
⭐ Drift detected. Do you retrain?Drift detection
In predictive maintenance, how do you know an alert was right?Drift detection
How would you check an LLM system for bias?Bias & explainability
⭐ Can you make it fair?Bias & explainability
⭐ Is chain-of-thought an explanation?Bias & explainability
A segment has 40 samples and looks bad.Bias & explainability

#Round 4 · Systems, serving and optimization

QuestionPage
⭐ Why does quantization make inference faster?Quantization
INT8 or INT4?Quantization
⭐ What breaks first when you quantize?Quantization
Why do large models quantize worse than small ones?Quantization
You got 1.2× not 4×. Why?Quantization
Quantization, pruning or distillation — which first?Distillation & pruning
⭐ You pruned to 90% sparsity and it is not faster.Distillation & pruning
Why do soft labels beat hard labels?Distillation & pruning
What is the risk in distilling from a commercial API?Distillation & pruning
What is continuous batching?Serving
⭐ Your p95 latency doubled. Debug it.Serving
⭐ How do you handle overload?Serving
How would you cut the bill in half?Serving
Why is max_tokens a capacity control?Serving

#Round 5 · Agents, safety and adaptation

QuestionPage
What is an agent?Agents
⭐ Why do long agent chains fail?Agents
How do you stop an agent running forever?Agents
⭐ How do you design tools?Agents
When is multi-agent worth it?Agents
How would you evaluate an agent?Agents
⭐ How do you prevent prompt injection?Guardrails
Direct versus indirect injection?Guardrails
Where do you enforce permissions?Guardrails
What is the dual-LLM pattern?Guardrails
⭐ RAG or fine-tuning?Fine-tuning
What is LoRA, mechanically?Fine-tuning
How much training data?Fine-tuning
⭐ When would you not fine-tune?Fine-tuning
What actually improves a prompt?Prompt engineering
How do you get reliable JSON?Prompt engineering
You have changed the prompt 15 times and it still fails.Prompt engineering

#Round 6 · Judgement, product and business

Asked at senior level, and where technical candidates most often stumble.

QuestionPage
⭐ How would you choose a model?Model selection
Hosted or self-hosted?Model selection
What licence questions matter?Model selection
Your context window is 200k. Use it?Model selection
LangChain or LlamaIndex?Orchestration
What is DSPy actually doing?Orchestration
⭐ Would you use a framework at all?Orchestration
⭐ Should we build or buy?Market & business
How would you price this?Market & business
⭐ What is our moat?Market & business
The demo works. Why isn't it shipped?Market & business
⭐ When should we not use an LLM?Market & business

#The twelve that matter most

If you have one evening, drill these. They cover the widest ground and each has a counter-intuitive core that most candidates miss.

  1. Why does quantization make inference faster? — bandwidth, not arithmetic
  2. The score went up 4 points on 50 items. Ship it? — no; ±11 point interval
  3. How do you know your judge is any good? — kappa against human labels
  4. How do you prevent prompt injection? — you don't; you bound the blast radius
  5. Why do long agent chains fail? — reliability compounds multiplicatively
  6. RAG or fine-tuning? — knowledge versus behaviour
  7. RAG quality is poor. Where first? — the retrieval ceiling
  8. Is chain-of-thought an explanation? — no, and saying why is the signal
  9. You pruned to 90% and it's not faster. — unstructured sparsity does nothing
  10. Can you make it fair? — not in all senses at once; it's a proven result
  11. What's our moat? — not the prompt, model or pipeline
  12. When should we not use an LLM? — when a deterministic solution exists

#Answering well, mechanically

Lead with the answer, then justify. "No — at n=50 the interval is ±11 points" beats three sentences of preamble arriving at the same place.

Name the trade-off. Almost every question here has one. Naming it unprompted is most of what "senior" sounds like.

Say what you would measure. The strongest ending to any answer is how you would know you were right.

Admit the limit. "I would need to check the current pricing" is stronger than a confident wrong number. Interviewers are calibrating whether they can trust what you say — a single fabricated detail costs more than the answer gains.

Use your own work. You have an eval harness with a golden set, a gate, and a failure-injection test you watched go red. Nearly every question in Rounds 3 and 4 can be answered with "here is what I actually did" — which beats any theoretical answer available to anyone else in the pipeline.