LLM Handbook

#Research papers

Most reading lists are undifferentiated: fifty papers, no guidance, so you read none of them. This one is sorted by what you should actually do with each paper, because that is the decision you are making.

MarkerMeans
📖 Read itSit down with it. The paper itself teaches something a summary cannot
📄 Skim itRead the abstract, figures and conclusion. Twenty minutes
🏷️ Name itKnow what it established and why. Reading it adds little

A caution before the list: papers report results under their own conditions. A method that beat baselines on MS MARCO in 2023 may do nothing for your corpus. Read them for mechanisms and for the shape of the trade-off, not for numbers to quote.


#If you only read ten

In this order. This is roughly two weekends and it covers the mechanisms underneath everything else.

#PaperWhy this one
1📖 Attention Is All You Need (2017)The architecture. Everything refers back to it
2📖 The Illustrated Transformer — Alammar (not a paper)Read alongside #1; the figures do what the prose cannot
3📖 Chain-of-Thought Prompting (2022)Where reasoning-in-context came from
4📖 Lost in the Middle (2023)Why long context is not the answer you hoped
5📖 ReAct (2022)The agent loop, stated plainly
6📖 Judging LLM-as-a-Judge (2023)Judge biases, and the validation step everyone skips
7📖 LoRA (2021)How adaptation became cheap
8📖 LLM.int8() (2022)Emergent outliers — why big models quantize badly
9📖 PagedAttention / vLLM (2023)Why modern serving looks the way it does
10📖 Precise Zero-Shot Dense Retrieval — HyDE (2022)Short, surprising, changes how you think about retrieval

Interactive simulation — needs JavaScript.


#Architecture & training

PaperYearWhat it establishedDo
Attention Is All You Need2017The transformer. Attention replaces recurrence📖
BERT2018Bidirectional pretraining; the encoder lineage🏷️
GPT-3 / Language Models are Few-Shot Learners2020In-context learning emerges with scale📄
Scaling Laws for Neural Language Models2020Loss follows predictable power laws in compute, data, parameters📄
Training Compute-Optimal LLMs (Chinchilla)2022Most models were badly undertrained on data📖
LLaMA2023Small models trained far longer beat larger undertrained ones📄
RoFormer (RoPE)2021Rotary position embeddings; relative position, extrapolable🏷️
GQA2023Grouped-query attention; shrinks the KV cache🏷️
Switch Transformer2021Sparse mixture-of-experts at scale🏷️

Chinchilla is the one to actually read in this section. It reframed the field — the finding that parameter count had been over-weighted relative to training tokens is why the useful small models exist at all.


#Prompting & reasoning

PaperYearWhat it establishedDo
Chain-of-Thought Prompting2022Intermediate reasoning steps improve multi-step tasks📖
Self-Consistency2022Sample several reasoning paths, take the majority answer📄
Tree of Thoughts2023Search over reasoning branches🏷️
Take a Step Back2023Abstracting the question first improves retrieval and reasoning📄
Calibrate Before Use2021Few-shot output is biased by example order and label distribution📖
Language Models Don't Always Say What They Think2023Stated reasoning can be unfaithful to the actual computation📖

The last one matters more than its citation count suggests. It is the evidence behind "chain-of-thought is not an explanation", and it is the paper to cite when someone proposes showing CoT to a regulator.


#Retrieval & RAG

PaperYearWhat it establishedDo
Dense Passage Retrieval (DPR)2020Dense retrieval beats BM25 with the right training📄
RAG (Lewis et al.)2020The name and the pattern🏷️
ColBERT2020Late interaction: per-token vectors, MaxSim scoring📖
ColBERTv22021Made it storage-practical📄
HyDE / Precise Zero-Shot Dense Retrieval2022Embed a hypothetical answer, not the question📖
Lost in the Middle2023Recall degrades in the middle of long contexts📖
Self-RAG2023Reflection tokens to critique retrieval and output📄
Corrective RAG (CRAG)2024Grade retrieval, correct before generating📄
RAPTOR2024Recursive summary trees over chunks📄
HNSW2016The graph index nearly everything uses📖
Product Quantization2011Vector compression for billion-scale search🏷️
Matryoshka Representation Learning2022Embeddings truncatable to fewer dimensions📄
BEIR2021Retrievers generalise worse across domains than benchmarks suggest📖

BEIR is the sobering one. It shows retrieval methods that dominate one benchmark falling behind BM25 on out-of-domain data — which is the honest argument for hybrid retrieval and for evaluating on your own corpus.


#Evaluation

PaperYearWhat it establishedDo
Judging LLM-as-a-Judge (MT-Bench, Chatbot Arena)2023Judge viability, and position/verbosity/self-enhancement bias📖
RAGAS2023Reference-free RAG metrics📄
HELM2022Multi-metric holistic evaluation, not a single score📄
Model Cards for Model Reporting2019Disaggregated reporting as a norm📖
Inherent Trade-Offs in Fair Classification2017Fairness criteria cannot all hold simultaneously📖
Learning under Concept Drift: A Review2019Drift taxonomy and detection methods📄

Chouldechova (2017) and Kleinberg et al. (2016) are the fairness impossibility results. Short, mathematical, and they settle an argument that otherwise recurs forever in product meetings.


#Efficiency & serving

PaperYearWhat it establishedDo
LLM.int8()2022Emergent outlier features above ~6.7B parameters📖
GPTQ2022Accurate one-shot 4-bit post-training quantization📄
AWQ2023Activation-aware weight quantization📄
SmoothQuant2022Shift difficulty from activations to weights📄
QLoRA2023Fine-tuning a 4-bit base with adapters📖
FlashAttention2022IO-aware attention; same maths, far less memory traffic📖
PagedAttention / vLLM2023KV cache paging; the modern serving design📖
Orca2022Iteration-level (continuous) batching📄
Speculative Decoding2022–23Draft-and-verify; faster with identical output📄
The Lottery Ticket Hypothesis2018Sparse subnetworks that train as well as the whole🏷️
SparseGPT2023One-shot pruning of large models📄
Are Sixteen Heads Really Better than One?2019Many attention heads are removable📄
The Unreasonable Ineffectiveness of the Deeper Layers2024Depth pruning often beats width pruning📄

#Agents, tools & adaptation

PaperYearWhat it establishedDo
ReAct2022Interleaved reasoning and acting📖
Reflexion2023Verbal self-critique between attempts📄
Toolformer2023Models learning when to call tools🏷️
LoRA2021Low-rank adaptation of frozen weights📖
InstructGPT / RLHF2022Preference alignment as a training stage📖
Direct Preference Optimization (DPO)2023Preference optimisation without a reward model📄
Constitutional AI2022AI feedback against written principles📄
LIMA20231,000 curated examples beat far larger noisy sets📖
Not What You've Signed Up For2023Indirect prompt injection via retrieved content📖

LIMA is the one to read if you are about to fine-tune. It is the strongest published argument that data curation, not volume, is the work.


#How to read one of these efficiently

Most papers do not need a linear read. A method that works:

  1. Abstract, then figures. The figures usually carry the contribution. If you cannot tell what the paper claims after the figures, the abstract was badly written, not you.
  2. Conclusion before the method. It tells you what to look for.
  3. Method section only if you need the mechanism. Often you do not.
  4. Skip the related work unless you are surveying the field.
  5. Read the limitations section. It is the most honest part of most papers and the fastest route to knowing whether it applies to you.
  6. Check the baselines. A large improvement over a weak baseline is a weak result, and this is where over-claiming hides.

Twenty minutes per paper on that method covers far more ground than two hours of linear reading, and retains more.


#Staying current without drowning

Papers are a poor way to track a fast field — by publication, results are months old. More useful:

SourceWhat it is good for
Model and API changelogsWhat actually changed in the thing you depend on
vLLM, llama.cpp, PEFT release notesWhat is now practical, not merely published
Anthropic / OpenAI engineering docsPrompting and tool-use guidance from people with the logs
Simon Willison's blogThe clearest sustained coverage of injection and practical LLM engineering
BEIR, MTEB, ANN-Benchmarks leaderboardsComparable numbers across methods
Conference proceedings (NeurIPS, ICML, ACL, EMNLP)Depth, once a year, when you have time

A discipline worth adopting: when a paper or release seems relevant, write one line in your own notes saying what you would change because of it. If you cannot write that line, you did not need the paper — and this is also the habit that turns reading into something you can talk about.


#An honest caveat

Citations here are given by title, primary author and year so you can find them. Verify details before quoting a specific number or author list — venues, author orders and reported figures are exactly the kind of detail that is easy to get subtly wrong, and being confidently wrong about a paper in an interview is worse than not having read it.

The safe form is: "the LLM.int8 paper showed that above roughly seven billion parameters, a small number of activation dimensions become extreme outliers, which is why naive INT8 breaks." That is the mechanism, it is what matters, and it does not depend on a number you half-remember.