LLM Handbook

#The LLM Handbook

A working engineer's handbook for building, testing, operating and arguing about LLM systems. Every topic is written at three depths, and every topic is examined from every seat that has to live with the decision — not just the coder's.

This handbook does not repeat what you already have. Where a subject is already covered well in the study folder (02_AI_CORE/03_genai_rag_agents.md covers transformers, RAG end-to-end, chunking, embeddings, reranking, agents, LangGraph, MCP, Text2SQL and guardrails), this handbook links there instead of writing it twice. What follows is the material that was missing.


#How each page is built

Every topic page follows the same eight-part shape, so you always know where to look for the thing you need:

PartWhat it gives you
1 · DiagramThe mental model as a picture, before any prose
2 · DesignThe components and why each one is load-bearing
3 · FlowThe sequence, in the order it actually happens
4 · UMLStructure: classes, states, or a sequence diagram
5 · ExampleCode you can run, not pseudocode
6 · DepthThe senior layer — failure modes, scale, trade-offs
7 · From each seatThe same topic seen by seven different roles
8 · Interview questionsWhat gets asked, and the answer sketch

Plus a stop condition: the sentence that tells you that you are done, so a topic has an end rather than dissolving into endless reading.

#The three depths

Every topic is layered, and the layers are labelled inline:

  • Basic — the definition and the one-sentence why. Enough to follow a conversation without nodding along blankly.
  • Intermediate — how to build it, what the knobs are, what breaks first. Enough to ship it.
  • Advanced — why the obvious approach is wrong at scale, what the research actually says, and where the trade-off genuinely bites.

#The seven seats

Section 7 of every page. The same technology looks completely different depending on what you are accountable for, and being able to switch seats mid-conversation is most of what "senior" means in an interview.

SeatThe question they are asking
UserDoes this help me, and can I tell when it is wrong?
CoderWhat do I type, and what will bite me at 2am?
TesterHow do I prove this works when the output is non-deterministic?
System designerWhat are the components, the failure modes, the budgets?
ArchitectWhat does this commit us to for the next three years?
CEOWhat does it cost, what does it earn, and what is the risk?
MarketWho else does this, what is commoditised, where is the moat?

Interactive simulation — needs JavaScript.


#Reading order

graph TD
  A[Start here] --> B[LLM foundations]
  B --> C[Retrieval and advanced RAG]
  C --> D[Evaluation and judging]
  D --> E[Inference optimization]
  E --> F[Orchestration frameworks]
  F --> G[Deployment and operations]
  G --> H[Market and business]
  D -.->|the gate everything else feeds| C

Evaluation is deliberately early. Every other decision in this handbook — which chunker, which reranker, whether the quantised model is good enough, whether the new prompt shipped an improvement — is unanswerable without it. It is also the single most common gap in an LLM engineer's interview answers.


#Status

All modules written. Nothing is padded to look finished.

ModulePages
LLM foundationsTransformers · Tokenization · Decoding · Model selection · Prompt engineering · Multimodal · Long context · Reasoning models
Retrieval & advanced RAGChunking · Embeddings & vector databases · Reranking · Query transformation & HyDE · Corrective & self-RAG · Text2SQL · GraphRAG
Evaluation & judgingLLM as a judge · Drift detection · Regression gates · Bias & explainability
Inference optimizationQuantization · Distillation & pruning
Orchestration frameworksLangChain, LlamaIndex, DSPy
Agents & safetyAgents & tool use · Guardrails & security
Training & adaptationFine-tuning · Synthetic data generation
Deployment & operationsServing & operations · Caching strategies
Market & businessUnit economics, build vs buy, where the moat is not
Practice & interviewsSystem design walkthroughs · Interview question bank · Learning paths
ReferenceGlossary · Research papers · Numbers to know · Anti-patterns

#A note on honesty

Two things this handbook will keep doing, because they are what separates a useful technical document from a confident one:

It names what it does not know. Where the research is contested, or a number depends on your workload, it says so rather than picking a convenient figure.

It shows the failure. Every technique here has a regime where it is the wrong choice, and that regime is stated as plainly as the benefits. A page that only tells you when something works has not taught you how to decide.