Part 5 ended on a block and a stack of twelve. GPT-2 is that stack plus five decisions around it. Its one checkpoint runs two different programs. Training scores a whole sequence in one call. Generation, what serving calls inference, writes one token at a time, in a loop, and most of what breaks in production breaks there.
I loaded OpenAI’s GPT-2 small weights into nanoGPT’s model.py again and ran both programs on the same weights. Then I added a cache to the loop, cast the weights to bf16, and compared what came out.
Five decisions around the stack
GPTConfig, model.py:108-116, fits on one screen: 1,024 positions, 12 blocks, 12 heads, width 768. The vocabulary field says 50,304, which the comment on that line calls padding to a multiple of 64, for speed. OpenAI’s checkpoint has 50,257 rows, and loading it sets the config to match. Around those numbers GPT-2 made five decisions, and three of them are already in this series.
The tokenizer is byte-level BPE, Part 1: every string encodes, so there is no unknown token. The normalisation is pre-LN, Part 5. The token table is tied to the output layer at model.py:138. One matrix serves both ends: the table Part 5 counted as 31% of the model. Untied, GPT-2 small would have 163 million parameters instead of 124 million.
The fourth is positions. GPT-2 learns them: wpe is a table of shape [1024, 768], one row per position, added to the token rows before the first block. A table has a last row. Ask forward for 1,100 positions and the assert at model.py:173 stops you:
AssertionError: Cannot forward sequence of length 1100, block size is only 1024The fifth is the initialisation. Each block makes two writes to the bus, Part 5’s name for the residual stream, so a stack of L blocks adds 2L of them. In nanoGPT, every matrix, the token and position tables included, starts as a random draw with standard deviation 0.02 (OpenAI’s own code used 0.01 for the position table); biases start at zero and LayerNorm at one. model.py:143-145 starts the two layers that make those writes at 0.02/√(2L) instead, so the sum doesn’t grow with depth. I measured it on untrained models. With the scaled init, the stream’s norm grew about 11 times through 12 blocks, and 11 times through 48. With the plain init it grew 55 times and 112 times. Training takes it from there.
Training is one call
Hand forward a sequence and its targets, the same sequence shifted by one, and it predicts every position at once. I gave it 92 positions of text. It returned logits of shape [1, 92, 50257]. Those are raw scores for every token at every position, and softmax turns them into probabilities. It also returned one loss, the cross-entropy at model.py:187: minus the log of the probability given to the right next token, averaged over all 92. Part 4’s mask is what makes this honest. Row t sees only rows 0..t, so no prediction can read its own answer.
That’s the program training runs at every step, for hundreds of thousands of steps. Every position is known in advance, so all of them go through the stack together, as a few large matrix multiplications. If you’ve written a batch job over a table, you know the shape of this program. On my laptop’s CPU, one forward over 1,024 positions took 0.23 seconds.
Without targets, forward takes a shortcut at model.py:190: it computes logits for the last position only. That line is the first sign of the second program.
Generation is a loop
To write text, the model has to see each token before it can predict the next. There is nothing to run in parallel across positions. Here is nanoGPT’s generate, model.py:312-328, condensed:
for _ in range(max_new_tokens): idx_cond = idx[:, -1024:] # crop to the position table logits, _ = self(idx_cond) # the whole sequence, again logits = logits[:, -1, :] / temperature # [B, 50257], last position only probs = F.softmax(logits, dim=-1) # top_k masking skipped here idx_next = torch.multinomial(probs, 1) # draw one token id idx = torch.cat((idx, idx_next), dim=1) # append it and go roundRead the self(idx_cond) line again. Every step runs the full forward over everything so far and keeps one row. If you’ve re-read a whole file to append one line to it, you’ve written this loop. With an 11-token prompt, 256 new tokens push 35,456 positions through the stack for 256 useful rows.
The fix is a KV cache. Part 4 gave every position a key, which later queries match against, and a value, which they average. Under Part 4’s mask, the key and value at position t depend only on positions 0..t, so once computed they never change. Keep them, one pair of tensors per block, and each step only pushes the new token through the stack and attends over the stored rows. Same 256 tokens: 266 positions. The price is memory. For GPT-2 small in fp32, that’s about 74 kB per position. nanoGPT doesn’t have one; I wrote one in about thirty lines on top of its modules.
The timer agrees, but not in proportion. Generating 64, 256 and 512 tokens took 2.9, 5.0 and 7.4 times longer without the cache, for 37, 133 and 261 times the positions. At batch 1, each step still reads all 124 million weights to produce one token, and the cache doesn’t shrink that. Writing 1,023 tokens with the cache took 9.3 seconds, about forty times the training forward over the same length.
The cache also breaks at the edge of decision four. Past 1,024 tokens, generate crops the left end. From the second block on, every stored key and value was computed while the dropped token was still there, and with learned positions every token also moves to a new row of wpe. The whole cache is stale. nanoGPT’s loop gets away with it because it recomputes everything anyway; a cached loop has to handle it.
- no cache, positions per sequence
- 35,456
- with cache
- 266
- fewer positions
- 133×
cache = 2 × 12 blocks × 266 positions × 768 × 4 bytes × 1
Drag the new tokens and watch the two counts part. Then pick XL in bf16 and raise the batch until the cache bar passes the weights.
Same weights, different bugs
Two programs over the same weights should agree. Mine did, almost. Over 200 greedy steps, always taking the highest-scoring token, both loops wrote the same text. Their logits differed by up to 0.011. Almost all of that was the same shift for every token, and softmax ignores it. The gap between the two best candidates moved by at most 0.00005; the narrowest gap at any step was 0.004, so the tokens held. A longer run or a bigger model can close that margin, and nothing will raise an error when it does.
The difference is floating point. After the prompt, every matrix multiplication inside the cached loop’s blocks sees one row instead of the whole sequence, so the library picks different kernels, the routines that do the arithmetic, and those add in a different order. Float addition isn’t associative, so the sums differ in the last bits.
Then change something your deployment changes routinely. I cast the same weights to bf16, Part 3’s 16-bit format, and ran the same greedy loop. The fp32 model continued the prompt with ” were not working.”; the bf16 one with ” were not being run on the same system”. They split at the third token and never came back together. Running the prompt eight times in one batch kept the tokens but moved the logits by up to 0.005 against the single-prompt run.
The usual fix people reach for is temperature 0. It doesn’t fix this, and in nanoGPT it isn’t even a value. generate divides the logits by the temperature, so 0 turns all 50,257 logits into infinities, softmax turns those into NaN, the float for an undefined result, and torch.multinomial raises. Most APIs that accept 0 treat it as a switch to argmax, picking the top score. That’s greedy decoding again. It’s only as deterministic as the weights, dtype, batch shape and kernels under it.
Be fair about the scale of this. Every run here was repeatable: the same seed on the same machine sampled the same 40 tokens twice. The drift appears when something underneath changes, and in a serving stack those things change without asking you.
Sampling is a policy on top
Greedy decoding is a choice, and GPT-2 small shows why people avoid it: within 200 tokens, the loop above was repeating its own sentence. Sampling draws from the distribution instead, and three knobs shape it before the draw.
Temperature divides the logits before softmax; below 1 it sharpens the distribution, above 1 it flattens it. Top-k keeps the k highest-scoring tokens and drops the rest. Top-p, also called nucleus sampling, keeps the smallest set whose probabilities add up to at least p. nanoGPT’s sample.py defaults to temperature 0.8 and top-k 200; it has no top-p.
What the knobs do depends on the prompt. After “The capital of France is”, GPT-2 small’s top candidate is ” the”, at 8.5%; ” Paris” is fifth, at 3.2%. At temperature 1, covering 90% of the probability takes 1,504 tokens. After “import numpy as”, ” np” alone has 85%, and 90% takes three. Entropy, measured in bits, puts a number on the spread: 8.7 for the capital prompt at temperature 1, as uncertain as a fair pick among about 400 tokens, and 1.5 for numpy, about three.
Top-k can’t tell those apart, so any single k you pick is wrong for one of them. At nanoGPT’s defaults, those 200 tokens hold 89% of the mass after the capital prompt. After the numpy prompt they keep 199 options that share under 5%. Top-p adapts: at the same temperature and p = 0.9, it keeps 237 tokens for one and a single token for the other.
The capital of France is▍
␣the8.5%␣now4.8%␣a4.6%␣France3.2%␣Paris3.2%␣in2.7%␣also2.6%␣not2.4%␣home2.3%␣still1.5%␣under1.4%␣located1.4%
- top-p keeps
- 1,504 tokens
- top-40 keeps
- 53% of the mass
- top-200 keeps
- 70% of the mass
- entropy
- 8.7 bits
At the starting settings, temperature 1 and p = 0.9, the capital prompt keeps 1,504 tokens and numpy three. Drag the temperature to 2 and watch numpy’s count grow. The top-40 and top-200 readouts show what a fixed k keeps.
What the ML engineer watches
Four habits, and the first is a test.
Test the cached loop against the full forward, on logits. Run both programs for a few hundred steps and compare the greedy tokens and, with a tolerance, the logits. Token equality alone can pass while the logits drift, as mine did.
Budget the cache like the weights. Per sequence it holds 2·L·T·d numbers for L blocks, T positions and width d: a key and a value for every block and position. GPT-2 small at 1,024 positions in fp32 is 75 MB. GPT-2 XL in bf16 is 315 MB per sequence, so about ten full-length sequences weigh as much as its 3.1 GB of weights. When you raise the batch size, you are making a memory decision, not only a throughput one.
Pin what you compare. Dtype, batch shape, library version and kernel all move logits. If your regression test diffs model outputs, fix all of them, or compare with a tolerance and log what changed.
Log the sampling settings with every output. Temperature, top-k, top-p and the seed are part of the program. Two runs that differ only in them behave differently, and an output logged without them can’t be reproduced.
None of this is modelling work either. The weights are one artifact. The code around them decides whether a position is scored or written, what survives between steps, and how a distribution becomes a token, and each of those is ordinary code you can test.