LLM Handbook

#System prompts

The one sentence: the "system prompt" you read about in a leak and the system prompt you get from an API differ by more than two orders of magnitude, because almost all of the former is product scaffolding and tool definitions that a raw API call does not include.

Prompt engineering covers the techniques that survive measurement. This page is about the artifact itself: what goes in it, in what order, how long it should be, and how to read the published and extracted ones without drawing the wrong conclusion.

Every figure below is a file size in bytes from a public corpus, chosen because it is the one property of these documents that is objectively checkable and reproducible. Divide by roughly 4 for a token estimate.


#1 · Diagram

  WHAT THE MODEL SEES BEFORE YOUR FIRST TOKEN

  ┌──────────────────────────────────────────────────────────┐
  │  1. identity / role        "you are ..."                 │  small
  │  2. behavioural rules      tone, refusals, formatting    │  small
  │  3. product context        date, user tier, surface      │  small
  │  4. TOOL DEFINITIONS       one schema per tool  ────────►│  LARGE
  │  5. injected context       retrieved docs, memory        │  varies
  └──────────────────────────────────────────────────────────┘
                              │
                              ▼
                    then, finally, your message


  THE SAME MODEL, THREE WAYS  (bytes)

    published prompt        18,813   ← behaviour only
    extracted, no tools     49,660   ← + product scaffolding
    extracted, with tools  179,561   ← + tool schemas

              tools are 72% of what the model actually reads

Interactive simulation — needs JavaScript.


#2 · Design — "system prompt" means two different things

This is the distinction that makes everything else make sense, and almost every online discussion of leaked prompts misses it.

Published promptExtracted prompt
ContainsBehavioural instructionsInstructions + tool schemas + product context
SourceThe vendor, dated and versionedReconstructed by a user, undated
VerifiableYesNo
Claude Opus 4.618,813 B179,561 B

They differ by 9.5×, and both are honestly labelled. The vendor publishes the part that describes behaviour. The extracted file also captures the tool definitions and product wiring that the deployed surface adds. Neither is lying; they are measuring different objects.

So when someone says "Claude's system prompt is 180 KB", the accurate statement is: the Claude.ai product sends roughly 180 KB, of which under 11% is the behavioural prompt Anthropic publishes.

#The part nobody budgets for: tool definitions

The corpus contains the controlled comparison — the same model with and without tools:

ModelNo toolsWith toolsTools' share
Claude Opus 4.649,660 B179,561 B72%
Claude Sonnet 4.647,414 B174,341 B73%

Two independent models, the same answer. Roughly three quarters of a deployed system prompt is tool schemas, which is the measured version of the claim in harness & loops that a tool definition is a prompt. Every tool you add is permanent context, paid on every request, before the user says anything.


#3 · The number that should change how you read leaks

API prompts and consumer-product prompts are not the same order of magnitude:

SurfaceBytes≈ tokens
gpt-5.5-pro-api835~210
gpt-5.5-api862~215
gpt-5.3-codex-api194~50
gpt-5.3-chat-api2,200~550
gpt-5.5-instant (product)85,051~21,000
gpt-5.5-thinking (product)116,100~29,000
gpt-5.6-sol (product)127,131~32,000

The API gives you a nearly empty system prompt. gpt-5.5-api at 862 bytes against gpt-5.5-thinking at 116,100 is a 135× difference, and the pattern holds across every API/product pair in the corpus. Everything you have read about a leaked ChatGPT or Claude.ai prompt — the personality rules, the formatting conventions, the tool suite, the safety scaffolding — is product engineering you do not inherit when you call the API.

That cuts both ways. You do not get their guardrails for free, and you are not fighting their instructions either. The behaviour you get from a raw API call is much closer to the model itself, which is why the same model can feel different through two surfaces.

The surface variation within one vendor is nearly as wide: claude-voice-mode is 1,085 B while claude-in-chrome is 74,050 B — a 68× spread driven almost entirely by how many tools the surface exposes.


#4 · How long should yours be?

The published Claude prompts form a clean time series on a consistent basis:

DateModelBytes
2024-07-12Claude Opus 32,154
2025-05-22Claude Opus 410,863
2026-02-05Claude Opus 4.618,813
2026-09-22Claude Opus 5.527,216

About 12.6× in twenty-six months, and the growth is not decoration — it is mostly accumulated edge cases, each one a behaviour someone had to specify after observing it go wrong.

Two things follow for your own prompt:

  • Starting long is the mistake. Those documents grew in response to observed failures. A 3,000-token prompt written before you have seen your system fail is guessing at which edge cases matter, and paying for the guess on every request.
  • Every line should have a story. If you cannot name the failure a sentence prevents, it is a candidate for deletion. That is the only reliable pressure against monotonic growth, because each individual addition always looks cheap.

The cost is real and compounding — see prompt engineering §6 on what a long prompt costs per request, and KV reuse on why a stable prefix is worth far more than a short one. A 20,000-token system prompt that never changes is cheaper than a 2,000-token one that varies per user, because the first is a cache hit and the second is not.

#Ordering

Put the invariant parts first and the variable parts last. Prefix caching matches on an exact prefix, so a date or a user name near the top invalidates everything after it. Identity → rules → tools → product context → retrieved content → user message, in that order, is both the conventional structure and the cache-optimal one.


#5 · Depth — reading an extracted prompt without being fooled

Extracted prompts are useful evidence and unreliable fact. The corpus makes both halves visible.

They are not reproducible. No capture date, no model build, no method. The repository's own README invites contributors to open a PR "if you got a different result", which is a fair admission that two extractions of the same model can disagree. Treat any single file as one observation, not a specification.

Implausible sizes are a signal. One entry in the corpus, claude-projects-thread-claude.md, is 533,064 bytes — roughly 133,000 tokens, most of a context window, for something labelled a system prompt. The likely explanation is that the capture includes conversation or project content rather than only the prompt. When a figure fails a sanity check, the extraction is the thing to doubt.

Prefer the dated ones. The corpus separates a versioned official/ folder from the extracted files, and where a vendor publishes its own prompts, those are strictly better: dated, attributable, and stable enough to cite.

What extracted prompts are genuinely good for is the thing no published prompt shows you: the shape of a deployed one. How tools are described, where refusal language sits, how much space product context takes, what gets repeated. Fifty of them side by side is real evidence about structure, and structure is robust to any individual file being slightly wrong.


#6 · From each seat

SeatWhat this means here
UserNever sees it, and it shapes everything — including why the same model feels different in two apps.
CoderYour tool schemas are ~72% of the prompt. Writing a tool description badly is the same mistake as writing the prompt badly, and costs the same tokens.
TesterThe prompt is an input. Version it, diff it, and re-run your evals when it changes — a prompt edit is a deploy.
System designerOrder by volatility: invariant first, per-request last. That single decision determines whether prefix caching works at all.
ArchitectCalling the API means building the scaffolding yourself. The 135× gap between an API prompt and a product prompt is the work the product team did that you now own.
CEOIt is paid on every request, forever. A 20k-token prompt at a million requests a day is a permanent line item, and the fix is deletion, not a discount.
Market"Leaked system prompt" headlines almost always describe product scaffolding, not a model. The interesting comparison is between surfaces, not vendors.

#7 · Interview questions

"How long should a system prompt be?" Shorter than you think, and it will grow anyway. Published Claude prompts went from ~2 KB to ~27 KB in two years, and that growth came from specified edge cases rather than design. Start with what you can justify, add only in response to an observed failure, and keep the order stable so it caches. A long stable prompt beats a short volatile one on cost.

"Does the API give me the same system prompt as the chat product?" No, and the gap is enormous — around 862 bytes against 116,000 for one recent OpenAI pair, roughly 135×. Consumer products ship extensive scaffolding: tools, personality, formatting, safety layers. Through the API you inherit none of it, which means neither their guardrails nor their constraints. That is the main reason the same model behaves differently through two surfaces.

"Where does most of a deployed system prompt go?" Tool definitions. The same Claude model measured with and without tools is 49,660 versus 179,561 bytes — tools are about 72% of it, and the figure reproduces on a second model. Adding a tool is a permanent context cost, paid before the user has said anything.

"How would you order a system prompt?" By volatility, invariant first. Identity, rules and tool schemas do not change; the date, the user and retrieved documents do. Prefix caching matches on an exact prefix, so one variable token near the top invalidates the cache for everything below it — which turns an ordering preference into a measurable bill.


#Stop condition

You are done with this page when you can:

  • State the difference between a published and an extracted system prompt, and why they differ by ~10×
  • Say what fraction of a deployed prompt is tool definitions, and why that matters for adding tools
  • Explain why an API call and a chat product behave differently with the same model
  • Order a system prompt by volatility, and say what that buys
  • Read a leaked prompt as evidence of structure rather than as a specification

#Sources worth reading

  • Anthropic's published system prompts — dated and versioned in the release notes. Where a vendor publishes its own, that is the citable source.
  • asgeirtj/system_prompts_leaks — the corpus every measurement on this page comes from. CC0. Useful as evidence of shape across ~50 systems; its extracted entries are undated reconstructions and should be read as observations, not specifications. Note that it separates a versioned official/ folder, which is the half to trust.
  • Prompt engineering for the techniques that go inside the document, and KV reuse for why a stable prefix is worth more than a short one.

Related: Harness & loops — the system prompt is one component of a harness, and tool schemas are another · Guardrails & security for why instructions in a prompt are not a security boundary · Caching for what a stable prefix is worth in money.