Part 4 ended on one stage. A decoder is a stack of blocks, twelve in GPT-2 small, and attention is one of two stages in each. This post reads a whole block the way you’d read a pipeline: what flows in, what each stage writes, and where the parameters and the cost end up.
They don’t end up in the same place.
I read Block in nanoGPT’s model.py with the shapes in the margin and loaded GPT-2’s weights into it. Then I broke it on purpose: I took blocks out one at a time, then made each one overwrite its input instead of adding to it.
The block is four lines
nanoGPT is a minimal GPT-2 in two files. Its README now points to a successor, but model.py is still the shortest honest read of a GPT-2 block. Block.forward is model.py:103-106:
def forward(self, x): # x: [B, T, 768], the stream x = x + self.attn(self.ln_1(x)) # read a normalised copy, add attention's output x = x + self.mlp(self.ln_2(x)) # read again, add the feed-forward's output return x # [B, T, 768]: the shape it came in withInside the two calls, model.py:52-92, six tensors are worth naming. The code below is condensed, without scaling, masking, or dropout. The shapes in the margin come from running a model with GPT-2 small’s shape, d = 768 and 12 heads:
qkv = self.c_attn(x) # [B, T, 2304] q, k, v side by sideq, k, v = qkv.split(768, dim=2) # three views, [B, T, 768] eachq = q.view(B, T, 12, 64).transpose(1, 2) # [B, 12, T, 64] 12 heads of 64; k, v alikeatt = (q @ k.transpose(-2, -1)).softmax(-1) # [B, 12, T, T] Part 4's squarey = (att @ v).transpose(1, 2).reshape(B, T, 768)y = self.c_proj(y) # [B, T, 768] attention outh = self.gelu(self.c_fc(x)) # [B, T, 3072] the feed-forward's hidden layerout = self.c_proj(h) # [B, T, 768] feed-forward outWith the stream itself, that’s the block: x, the projections, the scores, attention’s output, the hidden layer, and the feed-forward’s output. Three of them share the stream’s shape. Everything else is a view, an elementwise step, or the per-head [B, 12, T, 64] layout.
One division of labour is worth reading twice. Attention is the only stage that mixes positions: row t of its output is built from rows 0..t of the stream. LayerNorm and the feed-forward work on one position at a time.
LayerNorm subtracts each 768-number row’s mean and divides by its standard deviation, then applies a learned scale and shift; model.py:27 is one call to F.layer_norm. The feed-forward widens each row to 3,072, applies GELU (a smooth version of max(0, x)), and narrows it back. If you’ve written a map over rows, you’ve written the feed-forward’s control flow.
model.py writes it, with GPT-2 small's shape. The stream runs down the rail; each branch reads a normalised copy and adds its output back. The scores are priced as the eager path holds them; sdpa's fused kernels don't. Pick a tensor, then drag the sizes.- shape
- [1, 1,024, 2,304]
- made by
- c_attn · 768 × 2304
- parameters
- 1,771,776
- per token
- 3.5 M ops
- this batch, bf16
- 4.7 MB
Pick the feed-forward hidden layer and read its parameters. Then pick the scores: they have none. Now drag T and watch the attention term, the scores and the weighted sum. The parameters don’t move; the arithmetic does.
The plus sign is the architecture
Read the two lines again. Neither stage replaces x. Each reads a normalised copy and adds what it computed to the original. The stream that enters a block leaves it with two edits on top, and the next block reads the result.
That stream, [B, T, 768] from the embedding to the last block, is what the literature calls the residual stream; this post calls it the bus. Every stage reads it and writes to it; no stage owns it. If you’ve written total += delta in a loop, you know the contract: every stage adds to the running total, and none can take back what an earlier one added.
The numbers show the contract holding. I ran 1,024 tokens of an earlier post through GPT-2 small and measured lengths, a vector’s norm, as medians over positions. In blocks 1 to 10, each stage adds a vector 16% to 40% as long as the stream it reads. The stream that comes out points almost where it went in: the cosine similarity, where 1 means the same direction, is 0.93 to 0.97. The norm also grows as it goes, from 49 after block 0 to 480 at the end. Nothing on the bus is ever normalised in place. Each stage normalises its own copy, and ln_f normalises once more before the output layer turns the last row into logits, the next-token scores.
That was a decision. Call GPT-2’s order pre-LN: normalise inside the branch, x + attn(ln(x)), and leave the bus alone. The original transformer, and GPT-1, were post-LN: normalise after the add, x = ln(x + attn(x)), which rescales the bus itself at every stage. GPT-2’s paper records the move in one sentence, the same one that adds ln_f, and cites a precedent but gives no reason of its own.
A small test shows one. I trained two six-block models on this blog’s text, one character at a time, with the same settings and no warmup (warmup: starting with a small step size and ramping it up). Loss is the model’s average surprise at each next token; lower is better, and 0 would mean certain and right.
After 300 steps the pre-LN model reached about 2.4, better than the 2.56 a model gets from knowing only the previous letter. The post-LN one stalled at 3.2, roughly what letter frequencies alone give, and one seed run to 600 steps was still there. Given a 100-step warmup, it broke loose between steps 200 and 250. It’s a toy, with two seeds and losses read off training batches, but the direction matches the published analysis of why post-LN needs a warmup.
Back to GPT-2 small: I deleted blocks. The intact model scores a loss of 3.60 on those 1,024 tokens. Remove any one of blocks 1 to 10, so the stream passes straight through, and the loss comes out between 3.61 and 3.89. Remove block 0 and it jumps to 9.15.
Overwriting is worse than removing. Keep a block, but let its output replace the stream instead of adding to it. For blocks 1 to 8 the loss is above 9.6. For blocks 2 to 7 it is above 10.8, the score of a uniform guess over the vocabulary. Worse than guessing.
- block
- 5
- removed
- 3.83 (+0.23, +6.4%)
- overwriting
- 12.42 (+8.83, +246%)
- write length, vs the stream it reads
- attention 19% · feed-forward 22%
- cosine, in to out
- 0.97
Switch to overwriting and watch blocks 2 to 7 pass the dotted line. Then pick block 0.
Block 0 is the exception, and the norms say why. The embedding enters with a norm of 4.6, and block 0’s attention writes a vector seven times that length. After block 0 the stream is mostly block 0’s output, so it behaves less like an edit than like a second embedding. Overwriting there costs nothing, 3.59, because adding already amounted to replacing.
Two thirds of the weights
You can count one block’s parameters by hand, with d = 768 and the feed-forward width f = 4d. Attention’s two projections are 4d² + 4d: 2,362,368 numbers. The feed-forward’s two are 2·d·f + f + d, which is 8d² + 5d: 4,722,432. The two LayerNorms are 3,072. The feed-forward holds two thirds of every block in every GPT-2 size, because the ratio comes from the 4 * config.n_embd at model.py:82 and not from the width. Llama 3.1 8B changes both stages, and its feed-forward holds 81% of a block; Part 7 reads why.
Twelve blocks are 85 million parameters, and GPT-2 small has 124,439,808. The rest is almost all the token table: 38.6 million numbers, 31% of the model. It’s used twice, as Part 2’s lookup at the bottom and as the projection to logits at the top; model.py:138 ties them, the same array and not a copy. In GPT-2 XL the same table is 5.2% of the model. Small models spend nearly a third of their weights on the vocabulary; large ones barely notice it.
The cost is somewhere else
In software, a function’s size and its running time come from the same code. Here they come apart. The parameters sit mostly in the feed-forward. The arithmetic follows T.
A forward pass costs about two operations per token for each number in a weight matrix, a multiply and an add; Part 4 counted them as FLOPs. Call the count of those numbers N, so the parameters cost 2N per token. Add the attention term, 4·L·T·d with L the number of blocks, for the scores and the weighted sum over the full square. For GPT-2 small at 1,024 tokens that’s 247 million operations per token from the parameters and 38 million from attention. torch’s FlopCounterMode agrees when attention is written out, Part 4’s eager path: 284.8 million per token, against 285.1 from the formula.
On sdpa, the default, the same counter reported 247 million at both 256 and 1,024 tokens. It has no entry for the fused CPU attention kernel, the single routine sdpa calls instead of the four lines. The attention term never reaches its total, and nothing tells you so.
Inside one block, the attention term is 18% of the arithmetic at 1,024 tokens, and it passes the linear layers at 4,608 tokens, six times d. The clock agrees. On my laptop’s CPU, one block’s attention stage, projections included, took 38% of the block’s time at 256 tokens, 47% at 1,024, and 57% at 4,096, while holding a third of its parameters.
Be fair about the split. At the lengths GPT-2 was built for, the linear layers do most of the arithmetic, and the feed-forward is the larger half of that. The attention term only takes over past a length GPT-2 never accepts. Still, when you pick a model you pick two budgets. The parameters set the memory and are fixed at load; the context sets the cost and a growing share of the memory, and it arrives with each request.
What the ML engineer watches
Four habits, and the first two are arithmetic.
Count parameters by component, not in total. The first widget is one block’s count: 4d² for attention, 2·d·f for the feed-forward. Add the vocabulary times d, once, for the table. If you compare a 124M model with a 1.5B one, ask how much of each is vocabulary.
Estimate cost as 2N per token, and 6N to train. Training is a forward and a backward pass, and the backward, which computes the gradients, costs about two forwards. That’s 6N per token and 6·N·D for a run of D tokens. nanoGPT’s config for GPT-2, config/train_gpt2.py, runs 491,520 tokens per iteration for 600,000 iterations: 295 billion tokens. 6·N·D is 2.2 × 10²⁰ operations, 195 hours of one A100 at its bf16 peak, a day of an eight-GPU node. The file says to expect about five days and the README four, which puts the run at a fifth to a quarter of peak. nanoGPT’s own estimator adds the attention term, model.py:296, and at 1,024 tokens that term is another 15%.
Depth is latency. Spend the same 85 million block parameters four ways: 6 blocks of width 1,088, 12 of 768, 24 of 544, and 48 of 384. At batch 1, one token through the stack took 3.2, 4.0, 4.7, and 6.7 ms on my CPU. Each block adds a fixed overhead, and blocks run one after another. If you serve one request at a time, depth is what you wait for, even when the arithmetic is equal.
Measure a block before you drop it. Layer pruning (shipping with some blocks removed) and early exit (stopping at a middle block when the answer looks settled) lean on the bus, and the bus is uneven. Removing a middle block of GPT-2 small raises its loss by 0.5% to 8%; removing block 0 raises it by 155%, and the last block by 41%. The numbers above took a minute on a laptop, on text from this blog. Produce them on your own traffic.
None of this is modelling work either. A block is four lines around a sum, and the sum is what lets a stack of twelve behave like one pipeline. Where the weights sit and where the time goes are two different answers, and both are arithmetic you can do before loading anything.