Part two showed that a neural network is a stack of weighted sums. Sums need numbers. A sentence is not numbers. So before a language model can do anything with "The cat sat on the mat", the text has to become a list of numbers, and after the model has done its work, numbers have to become text again. This post covers both halves: tokens and embeddings.
Step one: chop the text into tokens
The obvious approach is one number per word. Give "the" the number 1, "cat" 2, and so on. It fails for two reasons. English has hundreds of thousands of words, plus names, typos, code and other languages, so the list never ends. And it treats "run", "running" and "runs" as three unrelated things.
The other extreme, one number per character, has a tiny vocabulary but makes every sequence very long, and the model has to learn from scratch that c-a-t means something.
Modern models use a middle path. A tokenizer splits text into tokens, which are chunks that are sometimes whole words, sometimes parts of words, sometimes punctuation. Common words are one token. Rare words are split into common pieces. "unbelievable" might become "un", "believ", "able". A typical vocabulary is somewhere between 30,000 and 200,000 tokens, and every one of them has an ID number.
"The cat sat on the mat"
→ ["The", " cat", " sat", " on", " the", " mat"]
→ [791, 8415, 7731, 389, 279, 2450]
Notice the leading spaces. In most tokenizers a space is glued to the word after it, so " cat" and "cat" are different tokens. That is why a model can behave slightly differently if your prompt has an extra space.
How the tokenizer decides where to cut
The splits are not designed by hand. The standard algorithm, byte pair encoding, starts with single characters and repeatedly merges the pair that occurs most often in a large sample of text. "t" and "h" become "th", "th" and "e" become "the", and so on, until the vocabulary hits its target size. The result reflects the training text: English words come out as single tokens, while a Chinese sentence or a chunk of rare code is cut into many more pieces.
That has practical consequences:
- Pricing and limits are in tokens. API prices are quoted per million tokens and the context window is measured in tokens. Roughly, a token is three quarters of an English word, so 1,000 tokens is about 750 words. Other languages often cost two or three times as many tokens for the same meaning.
- Models are bad at letters. Ask how many r's are in "strawberry" and a model may get it wrong. It never sees the letters. It sees perhaps two tokens, "str" and "awberry", and has to remember from training how those are spelled. Same reason they struggle to reverse a word or rhyme reliably.
- Numbers are chopped oddly. "12345" might be "123" and "45". Arithmetic on tokens like that is hard, which is part of why models do maths badly unless they write out the steps or call a calculator.
- JSON and code are token heavy. Every brace, quote and indentation space costs tokens. Minifying JSON before sending it to a model can cut the bill noticeably.
Step two: turn token IDs into meaning
Token ID 8415 is just a label. The number itself means nothing; 8416 is not "slightly more cat". The network needs something it can do arithmetic on, where similar words end up with similar numbers. That is an embedding.
An embedding is a list of numbers, a vector, one per token. Not one number but hundreds or thousands. A small model might use 768 numbers per token, a large one 12,000 or more. The model keeps a big table: one row per token in the vocabulary, each row a vector. Looking up a token means grabbing its row.
" cat" → [0.21, -0.83, 0.05, 1.12, ..., -0.44] (768 numbers)
" dog" → [0.19, -0.79, 0.11, 1.08, ..., -0.40]
" car" → [-0.65, 0.32, 0.88, -0.14, ..., 0.71]
Where do the numbers come from? They are weights. The embedding table is part of the network and it is trained by gradient descent exactly like every other weight. It starts random. During training, whenever two tokens tend to appear in the same kinds of places, the updates pull their vectors closer together, because that helps the model predict the next token. Nobody tells the model that cats and dogs are similar. It discovers that "cat" and "dog" both show up after "I fed the" and before "chased", and their vectors drift together.
Meaning as geometry
With every token as a point in a space of 768 dimensions, distance means similarity. Words for animals cluster. Words for countries cluster. Programming keywords cluster. You cannot picture 768 dimensions, but the maths works the same as in three.
The famous demonstration: take the vector for "king", subtract "man", add "woman", and the nearest token to the result is "queen". The direction from man to woman is roughly the same as the direction from king to queen. The model has encoded a relationship as a direction in space. The same trick works for Paris minus France plus Italy landing near Rome, or walking minus walk plus swim landing near swimming. It is not perfect and modern models do not rely on it directly, but it shows that the numbers carry real structure.
Individual dimensions rarely mean anything clean. There is no "dimension 47 is royalty". Meaning is spread across all of them together, which is why these representations are called distributed.
Embeddings outside language models
The same idea, turn a thing into a vector so that similar things are close, is used all over the place:
- Semantic search. Embed every document, embed the query, return the nearest documents. This finds "how do I reset my password" when the document says "changing your login credentials", which keyword search would miss.
- Retrieval augmented generation. Before asking an LLM a question about your own files, find the relevant chunks by embedding similarity and paste them into the prompt. This is how most "chat with your documents" products work.
- Recommendations. Embed users and products in the same space. Recommend products near the user.
- Duplicate detection and clustering. Near identical vectors mean near identical meaning, even if the wording differs.
Model providers sell embedding models as a separate, cheap API for exactly these uses. You send text, you get back a vector of a thousand or so floats, and you store it in a database that can do nearest neighbour lookups.
Where the position went
One thing is missing. The embedding for " cat" is the same wherever it appears, so "dog bites man" and "man bites dog" would look like the same bag of vectors. Word order matters, so the model adds a second vector to each token that encodes its position in the sequence: first, second, third. How that is done is an architecture detail, but the effect is that the network can tell which token came first.
The round trip
Put it together and the full pipeline for a language model is:
text → tokenizer → token IDs → embedding lookup → vectors
→ the network does its work →
vector for next token → compare to every row in the vocabulary
→ probability for each possible next token → pick one → token ID → text
The last steps run the first ones in reverse. The network produces one vector, scores it against every token in the vocabulary, and turns the scores into probabilities. "Paris" gets 0.92, "Lyon" gets 0.03, "the" gets 0.001. Then a token is chosen, appended to the input, and the whole thing runs again for the next one. How a token is chosen from those probabilities, and what "temperature" means, is covered in part five.
What comes next
We now have text as a sequence of vectors going into "the network does its work". For a modern language model that network is a transformer, and the thing that makes a transformer work is attention: a way for every token to look at every other token and decide which ones matter for predicting what comes next. That is part four.