How to Make Your AI Use Fewer Tokens, and What It Adds Up To Over a Year

By Sheng Pang · Published · 7 min read

If you use an AI coding agent or a chat model daily, you have a token bill, either in money or in a usage cap that runs out mid afternoon. Most people try to fix it by asking for shorter answers. That helps a little. The big wins are elsewhere, because the reply is the smallest part of what a model processes. This post covers where the tokens really go, what to write in your instructions to stop the waste, and a worked estimate of what it adds up to over a year.

Where the tokens go

A model has no memory between calls. Every message you send is packaged with the system prompt, your instruction file, the entire conversation so far and every file or tool result the agent has read, and the whole bundle is processed again. The reply comes back. Then on the next turn, all of that plus the reply plus your new message goes in again.

So in a typical agent session the input side, the stuff being re-read, is easily ten times the output side. Cutting the reply by half saves five percent. Cutting what gets re-read saves the other ninety five. Every technique below is about the input side, except one.

1. Keep the instruction file short

Your CLAUDE.md or AGENTS.md is sent on every single turn. A thousand word file is around 1,300 tokens per message, forever. Over a 40 turn session that is over 50,000 tokens on the file alone. Cut it to the things the agent gets wrong without it, move reference material to separate files the agent can open when needed, and use folder level files so area specific rules load only in that area. Our guide to instruction files covers what to keep.

2. Tell the agent how to read, not just what to do

An agent asked to fix a bug will happily read six whole files to find it. Each file read is thousands of input tokens that then sit in the context for the rest of the session. A few lines in the instruction file change this behaviour:

When looking for something, search with grep first and read only the
matching region, not whole files.

Do not read generated files, node_modules, lock files or anything in dist/.

Before reading a file over 300 lines, say why you need all of it.

This is the single biggest lever in agentic tools. Reading is where the tokens go, and agents read generously unless told otherwise.

3. Shape the reply

The one output side technique, and worth doing because output tokens cost several times more each. Models default to a preamble, a restatement of the task, the answer, a summary of the answer and an offer to do more. Instructions that remove the wrapping:

Answer first, no preamble. Do not restate my question or summarise what
you just said. Do not offer follow ups. Show a diff or the changed lines,
not the whole file, unless I ask.

"The changed lines, not the whole file" matters a lot in coding. Rewriting a 400 line file to change three lines is 400 lines of output tokens.

4. Keep the prefix stable so caching works

Providers offer prompt caching: if the beginning of your input is byte for byte identical to a recent request, that part is re-read at a large discount, commonly around a tenth of the normal input price. The catch is that it only works on an unchanged prefix. So put the stable material first, system prompt, instruction file, reference documents, and the changing material last. Do not put a timestamp or a random ID at the top of the prompt, it breaks the cache every turn. The agent tools do this for you as long as your instruction file is not changing mid session, so leave it alone once a session starts.

5. End sessions, do not extend them

Turn 60 of a session re-reads 59 turns of history, including every dead end and every file read that turned out to be irrelevant. When the task changes, start fresh. If you need to carry state, ask for a five line summary and paste it into the new session. Most agent tools have a compact command that does this in place; using it whenever the context passes about half full keeps every later turn cheap.

6. Send less data

Minify JSON before it goes in as input; the whitespace and indentation in a pretty printed document can be a third of its tokens, and our formatter does it in one click. Trim logs to the relevant lines. Strip base64 images and long hashes out of anything you paste. Write in English to the model where you can, since most other languages cost two to three times the tokens for the same content.

What it adds up to: a worked estimate

This is not a measured benchmark. Real usage varies too much for one number to be honest. It is a model of a plausible day with the assumptions written down, so you can swap in your own.

Assume a developer runs four agent sessions a day, each around 30 turns. Instruction file, history growth, file reads and reply length are the four things the techniques above change. Per session:

ComponentUntunedTunedWhat changed
Instruction file, re-sent 30 times1,300 × 30 = 39,000400 × 30 = 12,000Trimmed to a screen
File reads accumulated in context~60,000~20,000Grep first, partial reads, no generated files
History re-read across the session~250,000~120,000Smaller reads and replies compound, compact at midpoint
Replies (output)~15,000~8,000No wrapping, diffs not whole files
Total per session~364,000~160,000About 56 percent less

Then apply caching. In the untuned case assume little of the prefix is stable and 30 percent of input is served from cache. In the tuned case the instruction file and early context are stable and 70 percent is cached. Charging cached input at a tenth of the normal rate, the effective input cost drops by roughly a further half in the tuned case.

Put together across four sessions a day and about 230 working days a year:

UntunedTuned
Raw tokens per year~335 million~147 million
Effective billed input after caching~73 percent of raw~37 percent of raw
Effective cost, relative100~22

Under these assumptions the tuned setup costs about a fifth of the untuned one for the same work. Change the assumptions and the number moves, but not the shape: the savings come from re-reading less and caching more, and they compound because every turn carries everything before it. If you are on a subscription with a usage cap rather than paying per token, the same arithmetic shows up as the cap lasting a full day instead of running out after lunch.

What not to do

  • Do not starve the model of context it needs. An agent that cannot see the relevant code guesses, and a wrong guess costs more turns than the read would have. The goal is no waste, not no reading.
  • Do not turn off reasoning to save tokens on hard problems. Thinking tokens are expensive and they are the reason the answer is right. Save them on easy tasks, spend them on hard ones.
  • Do not compress your instructions into cryptic shorthand. A model reads terse, clear English fine. It reads your private abbreviations badly and does more turns to compensate.
  • Do not obsess over reply length. It is the small slice. Fix the reading first.

The five line version

Keep the instruction file under a screen. Tell the agent to grep before it reads and never read generated files. Ask for diffs, not whole files, and no wrapping around answers. Keep the prompt prefix stable so caching works. Start new sessions when the task changes. Do those and the bill drops by more than half without the model getting any dumber. For why the model is counting in tokens in the first place, see our post on tokens as the currency of AI.

← Back to all articles