Part 1
Inside the model
What a language model actually sees, and how it produces text.
01 Tokens
Models don't read letters or words. Text is split into tokens — common words whole, rare ones in pieces — by an algorithm that learned the pieces from frequency alone. Prices, limits and many odd behaviours are all measured in tokens.
A tiny training corpus: four words with their frequencies. Start by splitting every word into single characters. The vocabulary is just the alphabet.
02 Embeddings
Each token — and, with an embedding model, each whole passage — is turned into a long list of numbers. Things with similar meaning end up close together in that space, which is what makes semantic search possible.
a real embedding has hundreds of dimensions — this is a projection
An embedding turns a piece of text into a list of numbers — a point in a space with hundreds or thousands of dimensions. This is two of them, which is a lie, but a useful one.
03 Next-token prediction and sampling
At every step the model produces a probability for each possible next token, and one is drawn. Temperature and top-p shape that draw — which is why the same prompt can give different answers.
A language model never outputs a word directly. It scores every token in its vocabulary, and those scores are turned into a probability for each possible next token. Here are the top five.
04 Why models hallucinate
Trained to produce plausible text, a model will produce plausible text even when it has no reliable knowledge — in the same confident tone as everything else. Giving it the facts in context is the main fix.
For a fact that appears thousands of times in training data, the probability piles up on the right answer. The model is confident, and it is right.
Part 2
Talking to a model
The request, the budget, and getting answers you can use.
05 Context windows
The model has no memory between requests: everything it should know — instructions, history, documents — is sent every time, and must fit in its context window. Caching makes resending a long, stable prefix cheap.
Every request re-sends the whole conversation: system prompt, tool definitions, and all history. The window is a budget you refill from scratch on every turn.
06 Streaming responses
A long answer can take seconds to finish, but the first tokens exist almost immediately. Streaming sends them as they're generated, so users start reading in under a second.
A model generates one token at a time. A long answer can take many seconds to finish — but the first words exist almost immediately. Streaming is about showing them as they arrive.
07 Structured outputs
When code consumes the answer, ask for data, not prose. Give the API a schema, and generation is constrained to it — every response is valid JSON of exactly that shape.
Ask a model for JSON in plain words and you usually get it — wrapped in chatty prose, sometimes with a string where you wanted a number, occasionally with a missing brace. Fine for a demo; painful to parse at scale.
Part 3
Models that act
Giving a model tools, and letting it use them in a loop.
08 Tool calling
You describe functions with a name, a description and a schema. The model can't run them — it asks you to, by returning a structured tool call, and you send the result back.
tool definition (you send this)
{
name: "search_orders",
description: "Find orders for a customer.
Use when the user asks about an order's
status, contents, or delivery date.",
input_schema: {
type: "object",
properties: {
customer_id: { type: "string" },
status: { enum: ["open","shipped"] },
limit: { type: "integer" }
},
required: ["customer_id"],
additionalProperties: false
}
}model emits
validate
tool_result → back to the model
A tool is a name, a description, and a JSON Schema. The description is the part that decides whether the model reaches for it at all — it is prompt text, not documentation.
09 The agent loop
An agent is tool calling in a loop: the model calls a tool, reads the result, decides what to do next, and repeats until the task is done — or a limit you set stops it.
The loop starts with one user message. Everything the model will ever know about this task lives in that array — the API itself is stateless.
Part 4
Knowledge, quality, and safety
Making a model useful for your data — and knowing it works.
10 Retrieval (RAG)
For knowledge too big or too fresh for the prompt, split documents into chunks, find the ones relevant to each question by embedding similarity, and put just those in the context.
query
Can I return a $300 jacket after 3 weeks?
★ marks the chunk that actually contains the answer
RAG has one job: put the right few paragraphs in front of the model. The document is split into chunks ahead of time, and each chunk is embedded into a vector.
11 Prompting, RAG, or fine-tuning?
Three ways to adapt a general model: change what you ask, change what it can see, or change the model itself. They solve different problems — and the cheapest one is usually the right first step.
Prompting is instant: change the instructions, run it again. RAG needs a retrieval pipeline. Fine-tuning needs a dataset, a training run, and an evaluation before you know if it helped. Start with prompting, always.
12 Evals
Model output varies, so “it looked right when I tried it” isn't evidence. An eval is a fixed set of cases with a way to grade them, run on every change — the tests of LLM apps.
green = passes the grader · red = fails · grey = never tested
How most LLM features are tested: try one prompt, it looks right, ship it. This tells you the happy path works and nothing else.
13 Prompt injection
A model can't reliably tell your instructions apart from instructions hidden in the content it reads — a web page, an email, a document. Treat that content as untrusted and limit what the model is allowed to do with it.
You are a support agent. You may read tickets and send emails.
Summarise ticket #812 for me.
Customer reports the export button does nothing on Safari…
action the harness receives
A normal turn. Three blocks reach the model, and it has no structural way to tell them apart — they are all just text in one context window.