The T in GPT stands for transformer. It is the network design behind every major language model since 2018, and also behind most modern image, speech and protein models. Part three left us with a sentence turned into a list of vectors. This post explains what the transformer does with them, using as little maths as possible.
The problem: words mean different things in context
The embedding for "bank" is the same vector whether the sentence is about a river or about money. Yet the model needs to know which, or it cannot predict what comes next. So the first job of the network is to update each token's vector using the other tokens around it. After processing, the vector for "bank" in "I deposited money at the bank" should be pulled toward finance, and in "we sat on the bank of the river" toward geography.
Older designs did this by reading the sentence one token at a time, left to right, carrying a running summary. That has two problems. Information from early words fades by the time you reach the end of a long paragraph. And it is slow to train, because you cannot process word ten until you have processed word nine. Transformers fix both by letting every token look at every other token directly, all at once.
Attention in one picture
Take the sentence "The animal did not cross the street because it was too tired". What does "it" refer to? A human reads "tired" and knows "it" is the animal. If the sentence ended "too wide", "it" would be the street.
Attention lets the vector for "it" ask every other word in the sentence "how relevant are you to me?", get back a score for each, and then blend in the vectors of the high scoring words. In the "tired" sentence the score for "animal" comes out high and the updated vector for "it" now carries a lot of animal in it. In the "wide" sentence, "street" scores high instead. The same mechanism, different context, different result.
That is all attention is. For each token: score every other token, turn the scores into weights that add up to one, and take the weighted average of their vectors. The token now contains a summary of the parts of the sentence that matter to it.
Queries, keys and values
If you read anything about transformers you will meet the words query, key and value. They are the three roles a token plays in the scoring, and a filing cabinet is the usual analogy:
- The query is what a token is looking for. For "it", something like "I need a noun I could refer to".
- The key is what a token advertises about itself. For "animal", something like "I am a singular noun, a living thing".
- The value is what the token hands over if it is selected. The actual content that gets blended in.
Each of the three is produced from the token's vector by multiplying with a weight matrix, and those matrices are learned by gradient descent like everything else. The score between two tokens is how well the query of one matches the key of the other. Nobody programs "look for nouns". The model finds that matching queries to keys this way lowers its loss, and the behaviour emerges.
Many heads, many layers
One attention pass can only capture one kind of relationship at a time. So a transformer runs several in parallel, each with its own query, key and value matrices. These are attention heads. In a trained model, one head might track which pronoun refers to which noun, another might link a verb to its subject, another might attend to the previous token, another to matching brackets in code. A layer might have 32 or 96 heads and their outputs are stitched together.
Then comes a small ordinary neural network applied to each token separately, called the feed forward block. Attention moves information between tokens; the feed forward block processes it within each token. Researchers have found that much of a model's factual knowledge lives in these blocks.
Attention plus feed forward is one layer. Stack it 30 to 100 or more times. Each layer refines the vectors a little more. Early layers sort out grammar and local word sense, middle layers build up meaning across the sentence, late layers get ready to predict the next token. At the very end, the vector at the last position is scored against the vocabulary and becomes the probability distribution for the next token that part three described.
The causal mask: no peeking
A language model is trained to predict the next token. During training it sees whole documents, but if token five could attend to token six, predicting token six would be trivial and the model would learn nothing. So attention is masked: each token can only look at itself and the tokens before it. This is why these models are sometimes called causal or autoregressive. It is also why they generate left to right: each new token is produced from everything before it and nothing after.
Why the context window exists
Every token attends to every earlier token. For a prompt of 1,000 tokens that is roughly a million score computations per head per layer. For 100,000 tokens it is ten billion. Cost grows with the square of the length. That is the reason for the context window, the maximum number of tokens a model can handle at once. Bigger windows need enormous engineering effort, and even when the window is large, models tend to use information from the middle of a very long prompt less reliably than from the start and end.
It also explains a pricing quirk. Sending a 50 page document with every question is expensive, because the model re-reads the whole thing to produce each response. Providers offer prompt caching to reuse the computed vectors for a repeated prefix, which is why putting the fixed part of a prompt first saves money.
Why transformers won
The 2017 paper that introduced them was titled "Attention Is All You Need", and the title was the argument. Dropping the one token at a time design meant every token in a batch could be processed simultaneously, which is exactly what graphics cards are good at. Training got faster, so models got bigger, and it turned out that bigger transformers keep getting better in a predictable way as you add data and compute. That predictability, called scaling laws, is what convinced companies to spend billions on training. The architecture has not changed fundamentally since. Today's models are the 2017 design with refinements to the attention maths, better position encoding and a great deal of engineering.
Not just text
Because the transformer only sees a sequence of vectors, anything that can be cut into tokens can go in. Images are cut into patches, each patch embedded like a word. Audio is cut into short time slices. That is how one model can look at a screenshot and describe it: the image patches and the text tokens go through the same layers and attention lets the text attend to the picture.
What you now know
A transformer takes a sequence of token vectors and passes them through many layers. In each layer, attention lets every token gather information from the relevant earlier tokens, then a feed forward block processes each token. At the end, the final vector predicts a distribution over the next token. All the weights inside, embedding table, query, key and value matrices, feed forward blocks, are learned by the gradient descent from part two.
Train that on a few trillion tokens of internet text and you have a base model that continues any text you give it. It is not yet a chatbot. It will happily continue your question with three more questions, because that is what internet forums look like. Turning it into something that answers, follows instructions and declines to help with the wrong things is a separate process, covered in part five.