Tokens: The Strange Currency Every AI Model Runs On

By Sheng Pang · Published · 8 min read

Every AI model you talk to has a unit of account, and it is not the word. It is the token. Tokens are what the model reads, what it writes, what you are billed for, and what it runs out of when a long conversation suddenly forgets your name. Once you start seeing tokens instead of words, a lot of odd model behaviour stops being odd. Here is a tour.

A token is a chunk, not a word

Before a model sees your message, a small program called a tokenizer slices it into pieces from a fixed menu of somewhere between 30,000 and 200,000 chunks. Common English words are one chunk. Less common ones are two or three. Punctuation is its own chunk. A space usually gets glued to the front of the word that follows it.

"I love tokenizers"       → "I" " love" " token" "izers"        4 tokens
"Antidisestablishment"    → "Ant" "idis" "establish" "ment"     4 tokens
"the"                     → "the"                               1 token
" the"                    → " the"                              1 token, a different one

The menu was built by counting which pairs of characters appear together most often in a giant pile of text and merging them, over and over, until the menu was full. Nobody designed it. That is why the boundaries look arbitrary: they reflect what was common in the training data, not what a linguist would choose.

The rule of thumb for English is about three quarters of a word per token, so a thousand tokens is roughly 750 words, or a page and a half. Everything else in this post follows from the fact that the menu was learned from data that was mostly English prose and code.

The tax on every other language

Because the menu is full of English fragments, English is cheap. A sentence in Chinese, Japanese, Hindi or Arabic gets sliced much finer, often into single characters or even fragments of a character's byte encoding. The same meaning can cost two to three times the tokens. Newer tokenizers have improved this, but the gap has not closed.

This has real consequences. A non English user pays more per question, gets less useful history in the same context window, and hits the output limit sooner. If you are building a product for a non English market, the price per conversation is a different number than the one on the pricing page, and you should measure it.

Emoji, code and JSON are expensive too

Some more things that surprise people when they look at the token count:

  • A single emoji can be two to four tokens. The face with tears of joy costs more than the word "laughing".
  • Indentation is tokens. Code indented with eight spaces per level burns tokens on whitespace. A tokenizer built with code in mind has chunks for runs of spaces, but they still count.
  • JSON is quotes, braces, colons and repeated key names. A pretty printed JSON document can be a third bigger in tokens than the same document minified. The keys are repeated for every item in an array, so an array of a thousand objects with a key named "customerEmailAddress" spends a thousand times whatever that key costs.
  • Numbers are sliced into arbitrary chunks. "1234567" might be "123", "456", "7". The model then has to do arithmetic across chunk boundaries, which is one reason it is unreliable at maths.
  • Base64, hashes and UUIDs are worst of all. Random characters have no common pairs to merge, so a UUID can be twenty plus tokens for thirty six characters.

Paste a chunk of text into a tokenizer viewer once. Most providers publish one. Seeing your prose light up in coloured chunks is the fastest way to build an instinct for this.

Why the model cannot count the r's in strawberry

The most famous token failure. Ask a model how many r's are in "strawberry" and it may say two. It is not being careless. It never saw the letters. It saw perhaps "str" and "awberry", or "straw" and "berry", and the number of r's inside those chunks is something it has to remember from training rather than look at. Same story for reversing a word, finding the fifth letter, or writing a sentence where every word starts with the same letter. Anything at the level of individual characters is done blind.

The workaround is to ask the model to spell the word out with spaces first. "s t r a w b e r r y" makes each letter its own token, and now it can count. Newer models trained with this trick and reasoning models that think before answering get it right more often, but the underlying blindness is still there.

The model thinks in tokens too

Tokens are not only the input format. The model produces its answer one token at a time, choosing each from a probability distribution over the whole menu, then feeding it back in and choosing the next. There is no draft, no outline, no going back to fix the first sentence once it has seen the last. Whatever planning happens has to happen inside the computation that picks the next chunk.

This is why asking a model to "think step by step" works. Writing the steps out puts them into the token stream, where every later token can look back at them. The scratch work becomes part of the input. Reasoning models take this further and generate thousands of tokens of thinking before the first token of the answer, which is why they are slower and more expensive and better at hard problems.

Where your money goes

Pricing is per million tokens, with separate rates for input and output, and output usually costs several times more because generating is more work than reading. A typical conversation looks like this:

  • The system prompt. Hidden instructions the product sends before your message. In a coding agent this can be thousands of tokens of tool descriptions and rules, sent on every single turn.
  • The whole history. A model has no memory between calls. Every turn, the entire conversation so far is sent again as input. Turn 30 of a chat costs 30 turns' worth of input tokens. This is the quiet reason long conversations get expensive, and it is why providers offer prompt caching: if the first part of the input is identical to last time, they charge less to re-read it.
  • Tool results. When an agent reads a file or a web page, that content becomes input tokens. A single "read this 3,000 line file" is more tokens than an hour of chatting.
  • The reply. The expensive output tokens, and the only part most people think about.

Look at any real usage dashboard and the input side dwarfs the output side, often by ten to one. Which means the way to spend less is rarely "get shorter answers". It is "send less, and send the same prefix so it caches". We go into the practical side in how to set up instructions that cut token use.

The context window is a token budget

Every model has a maximum number of tokens it can hold at once, the context window. The reason it exists is that attention, the mechanism that lets each token look at the others, costs more as the square of the length. Windows have grown from a few thousand tokens to a million, but a bigger window is not free: long contexts are slower, cost more per turn, and models are measurably worse at using information from the middle of a huge prompt than from its start or end.

When a conversation overflows, products either refuse, silently drop the oldest turns, or summarise them. That is the moment a chatbot forgets what you told it an hour ago. It did not forget. The tokens were cut.

What to do with this knowledge

  • Write in English for the model when you can. It is cheaper and, for most models, slightly more accurate.
  • Minify JSON and strip whitespace before sending it as input. Our formatter has a minify button for exactly this.
  • Put the stable part of a prompt first so caching kicks in.
  • Start a fresh conversation when the topic changes. Old turns are dead weight you keep paying for.
  • For character level tasks, spell things out. For arithmetic, ask for the steps or give the model a calculator tool.
  • Do not ask a model how many tokens something is. It cannot see its own tokens either. Use the tokenizer tool.

Tokens are a leaky abstraction. They are an artefact of how the model was built, they show through in its behaviour, and they are the meter that runs while you use it. Learn to see them and both the bill and the weird failures make sense. If you want the deeper story of what happens to a token once it is inside the model, our AI basics post on tokens and embeddings picks up from here.

← Back to all articles