#Interview question bank
Every question from across the handbook, consolidated. Each links back to the page that answers it properly.
⭐ marks the questions that most reliably separate candidates — usually because the obvious answer is wrong, or because most people stop one level short.
#How to drill this
Do not read the answers. Cover them, answer out loud, then check.
The failure mode of a question bank is recognition without recall: you read the answer, it feels familiar, you conclude you know it, and then you cannot produce it under pressure. Say it out loud — the gap between "I know this" and "I can say this in sixty seconds" is where interviews are lost.
Time-box to two minutes per question. Most answers should take sixty to ninety seconds. If you cannot finish in two minutes you are including material that is not earning its place.
Interactive simulation — needs JavaScript.
#Round 1 · Foundations
Asked to check you understand the machine, not just the API.
| Question | Page |
|---|---|
| Explain attention. | Transformers |
| ⭐ Why is decode memory-bound but prefill compute-bound? | Transformers |
| What is the KV cache and why does it matter? | Transformers |
| Why decoder-only for generation? | Transformers |
| What does FlashAttention change? | Transformers |
| ⭐ Why can't the model count the letters in a word? | Tokenization |
| Why is non-English text more expensive? | Tokenization |
| Temperature 0 versus 0.7 — when each? | Decoding |
| ⭐ Is temperature 0 deterministic? | Decoding |
| top_k or top_p, and why? | Decoding |
| Why don't chat models use beam search? | Decoding |
| What is speculative decoding? | Decoding |
#Round 2 · RAG and retrieval
The largest section, because it is the most-asked area.
| Question | Page |
|---|---|
| How would you chunk a technical manual? | Chunking |
| ⭐ RAG quality is poor. Where do you look first? | Chunking |
| What is small-to-big retrieval? | Chunking |
| ⭐ How do you handle permissions in RAG? | Chunking |
| You are changing embedding model. What is involved? | Chunking |
| How does HNSW work? | Embeddings & vector DBs |
| ⭐ What recall does your ANN index actually get? | Embeddings & vector DBs |
| Is 0.82 cosine similarity good? | Embeddings & vector DBs |
| Which vector database, and why? | Embeddings & vector DBs |
| ⭐ Why two retrieval stages instead of one? | Reranking |
| How deep should the shortlist be? | Reranking |
| When does a reranker not help? | Reranking |
| ⭐ What is HyDE and why does it work? | Query transformation |
| HyDE hallucinates. Isn't that a problem? | Query transformation |
| Multi-query or reranking — which fixes what? | Query transformation |
| How do you fuse results from several queries? | Query transformation |
| How do you stop RAG answering from irrelevant context? | Corrective RAG |
| CRAG versus self-RAG? | Corrective RAG |
| ⭐ Your relevance grader is sometimes wrong. Is that a problem? | Corrective RAG |
#Round 3 · Evaluation
Where most candidates are weakest, and where a good answer is most noticeable.
| Question | Page |
|---|---|
| ⭐ How do you know your LLM judge is any good? | LLM as a judge |
| Pointwise or pairwise judging? | LLM as a judge |
| ⭐ The score went up 4 points on 50 items. Ship it? | LLM as a judge |
| The vendor updated the judge model. What now? | LLM as a judge |
| When would you not use an LLM judge? | LLM as a judge |
| How do you stop LLM quality regressing? | Regression gates |
| ⭐ Why gate per bucket as well as overall? | Regression gates |
| Should a green run update the baseline? | Regression gates |
| Someone edited the corpus. What should the gate do? | Regression gates |
| How do you know your gate works? | Regression gates |
| How would you detect drift in a live system? | Drift detection |
| Covariate versus concept drift? | Drift detection |
| ⭐ You have no labels. How do you monitor quality? | Drift detection |
| ⭐ Drift detected. Do you retrain? | Drift detection |
| In predictive maintenance, how do you know an alert was right? | Drift detection |
| How would you check an LLM system for bias? | Bias & explainability |
| ⭐ Can you make it fair? | Bias & explainability |
| ⭐ Is chain-of-thought an explanation? | Bias & explainability |
| A segment has 40 samples and looks bad. | Bias & explainability |
#Round 4 · Systems, serving and optimization
| Question | Page |
|---|---|
| ⭐ Why does quantization make inference faster? | Quantization |
| INT8 or INT4? | Quantization |
| ⭐ What breaks first when you quantize? | Quantization |
| Why do large models quantize worse than small ones? | Quantization |
| You got 1.2× not 4×. Why? | Quantization |
| Quantization, pruning or distillation — which first? | Distillation & pruning |
| ⭐ You pruned to 90% sparsity and it is not faster. | Distillation & pruning |
| Why do soft labels beat hard labels? | Distillation & pruning |
| What is the risk in distilling from a commercial API? | Distillation & pruning |
| What is continuous batching? | Serving |
| ⭐ Your p95 latency doubled. Debug it. | Serving |
| ⭐ How do you handle overload? | Serving |
| How would you cut the bill in half? | Serving |
Why is max_tokens a capacity control? | Serving |
#Round 5 · Agents, safety and adaptation
| Question | Page |
|---|---|
| What is an agent? | Agents |
| ⭐ Why do long agent chains fail? | Agents |
| How do you stop an agent running forever? | Agents |
| ⭐ How do you design tools? | Agents |
| When is multi-agent worth it? | Agents |
| How would you evaluate an agent? | Agents |
| ⭐ How do you prevent prompt injection? | Guardrails |
| Direct versus indirect injection? | Guardrails |
| Where do you enforce permissions? | Guardrails |
| What is the dual-LLM pattern? | Guardrails |
| ⭐ RAG or fine-tuning? | Fine-tuning |
| What is LoRA, mechanically? | Fine-tuning |
| How much training data? | Fine-tuning |
| ⭐ When would you not fine-tune? | Fine-tuning |
| What actually improves a prompt? | Prompt engineering |
| How do you get reliable JSON? | Prompt engineering |
| You have changed the prompt 15 times and it still fails. | Prompt engineering |
#Round 6 · Judgement, product and business
Asked at senior level, and where technical candidates most often stumble.
| Question | Page |
|---|---|
| ⭐ How would you choose a model? | Model selection |
| Hosted or self-hosted? | Model selection |
| What licence questions matter? | Model selection |
| Your context window is 200k. Use it? | Model selection |
| LangChain or LlamaIndex? | Orchestration |
| What is DSPy actually doing? | Orchestration |
| ⭐ Would you use a framework at all? | Orchestration |
| ⭐ Should we build or buy? | Market & business |
| How would you price this? | Market & business |
| ⭐ What is our moat? | Market & business |
| The demo works. Why isn't it shipped? | Market & business |
| ⭐ When should we not use an LLM? | Market & business |
#The twelve that matter most
If you have one evening, drill these. They cover the widest ground and each has a counter-intuitive core that most candidates miss.
- Why does quantization make inference faster? — bandwidth, not arithmetic
- The score went up 4 points on 50 items. Ship it? — no; ±11 point interval
- How do you know your judge is any good? — kappa against human labels
- How do you prevent prompt injection? — you don't; you bound the blast radius
- Why do long agent chains fail? — reliability compounds multiplicatively
- RAG or fine-tuning? — knowledge versus behaviour
- RAG quality is poor. Where first? — the retrieval ceiling
- Is chain-of-thought an explanation? — no, and saying why is the signal
- You pruned to 90% and it's not faster. — unstructured sparsity does nothing
- Can you make it fair? — not in all senses at once; it's a proven result
- What's our moat? — not the prompt, model or pipeline
- When should we not use an LLM? — when a deterministic solution exists
#Answering well, mechanically
Lead with the answer, then justify. "No — at n=50 the interval is ±11 points" beats three sentences of preamble arriving at the same place.
Name the trade-off. Almost every question here has one. Naming it unprompted is most of what "senior" sounds like.
Say what you would measure. The strongest ending to any answer is how you would know you were right.
Admit the limit. "I would need to check the current pricing" is stronger than a confident wrong number. Interviewers are calibrating whether they can trust what you say — a single fabricated detail costs more than the answer gains.
Use your own work. You have an eval harness with a golden set, a gate, and a failure-injection test you watched go red. Nearly every question in Rounds 3 and 4 can be answered with "here is what I actually did" — which beats any theoretical answer available to anyone else in the pipeline.