A language model reads and generates tokens.

Those pieces of text count toward API usage and context limits, so word counts alone cannot tell you what a request will cost or whether it will fit. Here is how tokenization affects those decisions.

A token is a chunk, not a word

Before a model sees your message, its tokenizer splits the text into pieces from a fixed vocabulary.

Common English words are one chunk. Less common ones are two or three.

Punctuation is its own chunk.

A space usually gets glued to the front of the word that follows it.

The following splits use cl100k_base, measured with tiktoken. Other encodings can split the same text differently.

"I love tokenizers"       → "I" " love" " token" "izers"        4 tokens
"Antidisestablishment"    → "Ant" "idis" "establish" "ment"     4 tokens
"the"                     → "the"                               1 token
" the"                    → " the"                              1 token, a different one

The menu was built by counting which pairs of characters appear together most often in a giant pile of text and merging them, over and over, until the menu was full.

Nobody designed it. That is why the boundaries look arbitrary: they reflect what was common in the training data, not what a linguist would choose.

The rule of thumb for English is about three quarters of a word per token, so a thousand tokens is roughly 750 words, or a page and a half.

Everything else in this post follows from the fact that the menu was learned from data that was mostly English prose and code.

Language changes token counts

Tokenization depends on the model and language.

Equivalent sentences can have different token counts, and different tokenizers can reverse the size difference. Count the exact translations you intend to send instead of applying one multiplier to every language.

If a translation uses more tokens with your chosen tokenizer, it consumes more of the context window and can cost more at the same token rate.

Another language or tokenizer may produce the opposite result. Measure representative conversations for your product rather than applying an English word estimate.

Emoji, code and JSON are expensive too

Some more things that surprise people when they look at the token count:

Single examples depend on the encoding. With cl100k_base, 😂 uses two tokens and laughing uses three. With o200k_base, they use one and two respectively.

Indentation is tokens. Code indented with eight spaces per level burns tokens on whitespace. A tokenizer built with code in mind has chunks for runs of spaces, but they still count.

JSON is quotes, braces, colons and repeated key names. Count your own formatted and compact documents; whitespace does not imply a fixed saving. The keys are repeated for every item in an array, so an array of a thousand objects with a key named "customerEmailAddress" spends a thousand times whatever that key costs.

Numbers are sliced into arbitrary chunks. "1234567" might be "123", "456", "7". The model then has to do arithmetic across chunk boundaries, which is one reason it is unreliable at maths.

Base64, hashes and UUIDs can use many tokens. Random characters have no common pairs to merge, so a UUID can be twenty plus tokens for thirty six characters.

Paste a chunk of text into a tokenizer viewer once.

Most providers publish one. You can compare the displayed chunks with the words you typed.

Why the model cannot count the r's in strawberry

A character counting example.

Ask a model how many r's are in "strawberry" and it may say two. It is not being careless.

It never saw the letters.

cl100k_base splits strawberry into str, aw and berry. o200k_base produces st, raw and berry.

Same story for reversing a word, finding the fifth letter, or writing a sentence where every word starts with the same letter.

Anything at the level of individual characters is done blind.

The workaround is to ask the model to spell the word out with spaces first.

Spacing the letters makes them explicit in the prompt. Verify the result; this changes the input but does not guarantee a correct count.

The model thinks in tokens too

Tokens are not only the input format.

The model produces its answer one token at a time, choosing each from a probability distribution over the whole menu, then feeding it back in and choosing the next. There is no draft, no outline, no going back to fix the first sentence once it has seen the last.

Whatever planning happens has to happen inside the computation that picks the next chunk.

This is why asking a model to "think step by step" works. Writing the steps out puts them into the token stream, where every later token can look back at them.

The scratch work becomes part of the input.

Reasoning models take this further and generate thousands of tokens of thinking before the first token of the answer, which is why they are slower and more expensive and better at hard problems.

Where your money goes

Pricing is per million tokens, with separate rates for input and output, and output usually costs several times more because generating is more work than reading.

A typical conversation looks like this:

The system prompt. Hidden instructions the product sends before your message. Their loading and caching depend on the product.

The whole history. The product chooses which earlier messages and results to send. It may trim or summarize them, and prompt caching can change the price of repeated input. Check actual usage rather than assuming every turn resends everything.

Tool results. When an agent reads a file or a web page, that content becomes input tokens. Read relevant content while keeping the context needed for the task.

The reply. The generated answer is output, priced separately from input.

Compare input, cached input and output in your own usage report.

Their relative sizes depend on the task. Use that split to choose between trimming data, caching repeated context and shortening replies.

Our instruction and token usage guide gives changes to test.

The context window is a token budget

Every model has a maximum number of tokens it can hold at once, the context window. The reason it exists is that attention, the mechanism that lets each token look at the others, costs more as the square of the length.

A larger window gives you room for more input. Measure cost, latency and retrieval accuracy on your own prompts rather than treating window size as a quality guarantee.

When a conversation overflows, products either refuse, silently drop the oldest turns, or summarise them. That is the moment a chatbot forgets what you told it an hour ago.

It did not forget.

The tokens were cut.

What to do with this knowledge

Use the language the task needs. Compare exact token counts if you are choosing between equivalent prompts in different languages.

Minify JSON and strip whitespace before sending it as input. Our formatter has a minify button for exactly this.

Put the stable part of a prompt first so caching kicks in.

Start a fresh conversation when the topic changes. Old turns are dead weight you keep paying for.

For character level tasks, spell things out. For arithmetic, ask for the steps or give the model a calculator tool.

Do not ask a model how many tokens something is. It cannot see its own tokens either. Use the tokenizer tool.

Tokens are a leaky abstraction.

They are an artefact of how the model was built, they show through in its behaviour, and they are the meter that runs while you use it. Learn to see them and both the bill and the weird failures make sense.

If you want the deeper story of what happens to a token once it is inside the model, our AI basics post on tokens and embeddings picks up from here.