Streaming LLM Responses: Show the First Word in Under a Second

October 7, 2026 · 3 min read

Language models write one token at a time. A detailed answer might take five or ten seconds to finish — but the first few words are ready in a fraction of a second. Without streaming, the user stares at a spinner for the whole time. With it, they start reading almost immediately.

The total time doesn't change. What changes is time to first token, and that's the latency people actually feel.

browser
your server
model API
what the user sees · 0.0 s
(waiting…)

A model generates one token at a time. A long answer can take many seconds to finish — but the first words exist almost immediately. Streaming is about showing them as they arrive.

0 / 7

How it works

A normal request returns one response at the end. A streaming request keeps the HTTP connection open, and the server sends a sequence of server-sent events (SSE) as generation happens. With the Anthropic API the events look like this:

event: message_start          → the message begins
event: content_block_start    → a new block (text, tool use, …)
event: content_block_delta    → a few more characters of text
event: content_block_delta    → …
event: content_block_stop
event: message_delta          → stop_reason and final usage
event: message_stop

You don't parse these by hand — the SDK turns them into an async iterator.

Streaming on the server

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

const stream = client.messages.stream({
  model: 'claude-opus-5-5',
  max_tokens: 64000,
  messages: [{ role: 'user', content: 'Explain DNS to a beginner.' }],
});

for await (const event of stream) {
  if (event.type === 'content_block_delta' && event.delta.type === 'text_delta') {
    process.stdout.write(event.delta.text);
  }
}

const message = await stream.finalMessage();   // the complete response
console.log(message.stop_reason, message.usage);

Streaming also protects you from HTTP timeouts on long generations — for large max_tokens values, the SDKs expect you to stream.

Relaying it to the browser

Keep the API key on your server, and forward chunks to the page. In a Next.js route handler, that can be a streamed Response:

// app/api/chat/route.ts
export async function POST(req: Request) {
  const { question } = await req.json();

  const stream = client.messages.stream({
    model: 'claude-opus-5-5',
    max_tokens: 64000,
    messages: [{ role: 'user', content: question }],
  });

  const encoder = new TextEncoder();
  const body = new ReadableStream({
    async start(controller) {
      for await (const event of stream) {
        if (event.type === 'content_block_delta' && event.delta.type === 'text_delta') {
          controller.enqueue(encoder.encode(event.delta.text));
        }
      }
      controller.close();
    },
    cancel() {
      stream.abort();   // the user left — stop generating (and paying)
    },
  });

  return new Response(body, { headers: { 'Content-Type': 'text/plain; charset=utf-8' } });
}

And reading it in the page:

const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ question }) });
const reader = res.body!.pipeThrough(new TextDecoderStream()).getReader();

let answer = '';
while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  answer += value;
  render(answer);
}

The details that matter

Check the stop reason. end_turn means the model finished. max_tokens means the answer was cut off — show that to the user rather than presenting half an answer as complete.

Let users stop. A Stop button that cancels the fetch with an AbortController should close the stream all the way back to the model API, so you stop paying for tokens nobody will read. The cancel() handler above does that.

Errors can arrive mid-stream. The response may have started with a 200 and still fail halfway. Handle a broken stream separately from a failed request, and decide whether to retry or show what arrived.

Render efficiently. Dozens of deltas a second can trigger dozens of re-renders. Batch updates to one per animation frame, and render Markdown incrementally rather than re-parsing everything each time.

Structured data streams too. Tool calls and JSON arrive as partial input deltas. Show progress if you like, but only act on them once the block is complete and validated.

When not to stream

  • Background jobs where nobody is watching (batch processing is cheaper).
  • Outputs you must validate as a whole before showing anything — for example a JSON object that drives UI.

The takeaway

Streaming doesn't make the model faster; it makes the wait visible and useful. Stream from your server, forward chunks to the browser, check the stop reason, and wire cancellation all the way through — and a ten-second answer starts feeling instant.