← All guides

How LLMs Work

Everything between “type a prompt” and “ship an AI feature”: what the model sees, how it picks words, and how to make it useful, reliable and safe. No machine-learning background needed — 13 short chapters, each with an animation.

Press Play on any animation, or use Next and Back to go at your own pace. Mark chapters as done to track where you are — progress is saved in this browser.

0 of 13 chapters donesaved in this browser only
  1. Part 1 · 4 chaptersInside the modelWhat a language model actually sees, and how it produces text.
  2. Part 2 · 3 chaptersTalking to a modelThe request, the budget, and getting answers you can use.
  3. Part 3 · 2 chaptersModels that actGiving a model tools, and letting it use them in a loop.
  4. Part 4 · 4 chaptersKnowledge, quality, and safetyMaking a model useful for your data — and knowing it works.

Part 1

Inside the model

What a language model actually sees, and how it produces text.

01 Tokens

Models don't read letters or words. Text is split into tokens — common words whole, rare ones in pieces — by an algorithm that learned the pieces from frequency alone. Prices, limits and many odd behaviours are all measured in tokens.

×5
low
×2
lower
×6
newest
×3
widest
merges, in order
none yet

A tiny training corpus: four words with their frequencies. Start by splitting every word into single characters. The vocabulary is just the alphabet.

0 / 6
Key idea: The model sees a sequence of token IDs, not characters. Count tokens, not words.

02 Embeddings

Each token — and, with an embedding model, each whole passage — is turned into a long list of numbers. Things with similar meaning end up close together in that space, which is what makes semantic search possible.

dogpuppycatkittenbank (river)loanmortgagebank (money)

a real embedding has hundreds of dimensions — this is a projection

An embedding turns a piece of text into a list of numbers — a point in a space with hundreds or thousands of dimensions. This is two of them, which is a lie, but a useful one.

0 / 6
Key idea: Meaning becomes position: similar text, nearby vectors.

03 Next-token prediction and sampling

At every step the model produces a probability for each possible next token, and one is drawn. Temperature and top-p shape that draw — which is why the same prompt can give different answers.

The cat sat on the ▍
raw scores (logits) → probabilities at temperature 1
mat
52%
floor
23%
sofa
16%
roof
7%
moon
2%

A language model never outputs a word directly. It scores every token in its vocabulary, and those scores are turned into a probability for each possible next token. Here are the top five.

0 / 5
Key idea: A model outputs probabilities; sampling picks the words. Variation is by design.

04 Why models hallucinate

Trained to produce plausible text, a model will produce plausible text even when it has no reliable knowledge — in the same confident tone as everything else. Giving it the facts in context is the main fix.

The capital of Australia is ▍
a well-known fact
Canberra
91%
Sydney
6%
Melbourne
2%
a
1%

For a fact that appears thousands of times in training data, the probability piles up on the right answer. The model is confident, and it is right.

0 / 3
Key idea: Plausible isn’t the same as true. Ground answers in sources the model can copy from.

Part 2

Talking to a model

The request, the budget, and getting answers you can use.

05 Context windows

The model has no memory between requests: everything it should know — instructions, history, documents — is sent every time, and must fit in its context window. Caching makes resending a long, stable prefix cheap.

context window8,800 / 200,000 tokens
system 2,000tools 6,000first message 800 

Every request re-sends the whole conversation: system prompt, tool definitions, and all history. The window is a budget you refill from scratch on every turn.

0 / 7
Key idea: Each request starts from zero. Budget the window like memory, and keep stable parts first.

06 Streaming responses

A long answer can take seconds to finish, but the first tokens exist almost immediately. Streaming sends them as they're generated, so users start reading in under a second.

browser
your server
model API
what the user sees · 0.0 s
(waiting…)

A model generates one token at a time. A long answer can take many seconds to finish — but the first words exist almost immediately. Streaming is about showing them as they arrive.

0 / 7
Key idea: Streaming doesn’t make the model faster — it makes the wait useful.

07 Structured outputs

When code consumes the answer, ask for data, not prose. Give the API a schema, and generation is constrained to it — every response is valid JSON of exactly that shape.

const Invoice = z.object({
vendor: z.string(),
total: z.number(),
dueDate: z.string(),
})
 
const res = await client.messages.parse({
model: 'claude-opus-5-5',
max_tokens: 1024,
messages: [{ role: 'user', content: email }],
output_config: { format: zodOutputFormat(Invoice) },
})
const invoice = res.parsed_output
asking for JSON in the prompt only
reply
Sure! Here's the invoice data: ```json { "vendor": "Acme", "total": "1,240.00", … } ``` Let me know if…

Ask a model for JSON in plain words and you usually get it — wrapped in chatty prose, sometimes with a string where you wanted a number, occasionally with a missing brace. Fine for a demo; painful to parse at scale.

0 / 4
Key idea: Schemas guarantee the shape. You still have to check the values.

Part 3

Models that act

Giving a model tools, and letting it use them in a loop.

08 Tool calling

You describe functions with a name, a description and a schema. The model can't run them — it asks you to, by returning a structured tool call, and you send the result back.

tool definition (you send this)

{
  name: "search_orders",
  description: "Find orders for a customer.
    Use when the user asks about an order's
    status, contents, or delivery date.",
  input_schema: {
    type: "object",
    properties: {
      customer_id: { type: "string" },
      status: { enum: ["open","shipped"] },
      limit: { type: "integer" }
    },
    required: ["customer_id"],
    additionalProperties: false
  }
}

model emits

waiting…

validate

—

tool_result → back to the model

—

A tool is a name, a description, and a JSON Schema. The description is the part that decides whether the model reaches for it at all — it is prompt text, not documentation.

0 / 7
Key idea: The model chooses and fills in the call; your code runs it.

09 The agent loop

An agent is tool calling in a loop: the model calls a tool, reads the result, decides what to do next, and repeats until the task is done — or a limit you set stops it.

call model
→read stop_reason
→run tools
→append results
turn 1
"How many open PRs need review?"user · text
stop_reason:—1,240 input tokens resent

The loop starts with one user message. Everything the model will ever know about this task lives in that array — the API itself is stateless.

0 / 6
Key idea: Call, observe, decide, repeat — with a budget and a stopping rule.

Part 4

Knowledge, quality, and safety

Making a model useful for your data — and knowing it works.

10 Retrieval (RAG)

For knowledge too big or too fresh for the prompt, split documents into chunks, find the ones relevant to each question by embedding similarity, and put just those in the context.

query

Can I return a $300 jacket after 3 weeks?

c1Refund policy — customers may return items within 30 days…—
c2…for orders over $200, returns require a support approval…—
c3Shipping: standard delivery takes 3–5 business days…—
c4Warranty claims are handled by the manufacturer, not us…—
c5Refunds are issued to the original payment method within…—
answer —

★ marks the chunk that actually contains the answer

RAG has one job: put the right few paragraphs in front of the model. The document is split into chunks ahead of time, and each chunk is embedded into a vector.

0 / 5
Key idea: If the right chunk isn’t retrieved, the model can’t use it — retrieval quality is the job.

11 Prompting, RAG, or fine-tuning?

Three ways to adapt a general model: change what you ask, change what it can see, or change the model itself. They solve different problems — and the cheapest one is usually the right first step.

need 1 of 6
Try an idea this afternoon
Promptingbest fit
RAGworkable
Fine-tuningpoor fit

Prompting is instant: change the instructions, run it again. RAG needs a retrieval pipeline. Fine-tuning needs a dataset, a training run, and an evaluation before you know if it helped. Start with prompting, always.

0 / 5
Key idea: Prompt first. Retrieve facts. Fine-tune behaviour, not knowledge.

12 Evals

Model output varies, so “it looked right when I tried it” isn't evidence. An eval is a fixed set of cases with a way to grade them, run on every change — the tests of LLM apps.

refund, in policy
refund, expired
order status
multi-item order
ambiguous question
no data in corpus
prompt injection
non-English
score
v1
1/1

green = passes the grader · red = fails · grey = never tested

How most LLM features are tested: try one prompt, it looks right, ship it. This tells you the happy path works and nothing else.

0 / 6
Key idea: You can’t improve what you don’t measure — and one sample isn’t a measurement.

13 Prompt injection

A model can't reliably tell your instructions apart from instructions hidden in the content it reads — a web page, an email, a document. Treat that content as untrusted and limit what the model is allowed to do with it.

system prompttrusted · yours

You are a support agent. You may read tickets and send emails.

user messagesemi-trusted · the user

Summarise ticket #812 for me.

tool_result · ticket #812untrusted · fetched content

Customer reports the export button does nothing on Safari…

action the harness receives

executed — summarise ticket

A normal turn. Three blocks reach the model, and it has no structural way to tell them apart — they are all just text in one context window.

0 / 6
Key idea: Anything the model reads can try to give it orders. Design permissions as if it will succeed.
That's the tour. To build on it, see the AI Engineering articles, or try the system design practice tool, which calls a model straight from your browser with your own key. Code in the linked articles uses @anthropic-ai/sdk.