A coding agent can spend tokens reading files, carrying conversation history and producing replies. Shorter answers address only one part of that usage.

Compare the input and output totals for your own sessions before choosing what to trim.

Where the tokens go

A model call receives the context the product sends. A coding agent may include instructions, earlier messages, files and tool results.

Products can trim, summarize or cache that material, so inspect your own usage report before assuming that each turn processes the entire history again.

The split between input, cached input, reasoning and visible output depends on the task and product. Use your usage report to identify the expensive part.

The techniques below give you changes to test, rather than a universal saving percentage.

1. Keep the instruction file short

Instruction files can contribute to the context a coding agent loads. The exact files and loading behavior depend on the tool.

Keep required rules concise, move reference material into separate files when appropriate, and compare measured usage.

Our instruction file guide covers the choices.

2. Tell the agent how to read, not just what to do

An agent asked to fix a bug will happily read six whole files to find it.

Each file read is thousands of input tokens that then sit in the context for the rest of the session. A few lines in the instruction file change this behaviour:

When looking for something, search with grep first and read only the
matching region, not whole files.

Do not read generated files, node_modules, lock files or anything in dist/.

Before reading a file over 300 lines, say why you need all of it.

Searching before reading can reduce irrelevant file content. Preserve the context needed to understand a bug; a partial read that hides a dependency can create extra work.

3. Shape the reply

The one output side technique, and worth doing because output tokens cost several times more each.

Models default to a preamble, a restatement of the task, the answer, a summary of the answer and an offer to do more. Instructions that remove the wrapping:

Answer first, no preamble. Do not restate my question or summarise what
you just said. Do not offer follow ups. Show a diff or the changed lines,
not the whole file, unless I ask.

"The changed lines, not the whole file" matters a lot in coding. Rewriting a 400 line file to change three lines is 400 lines of output tokens.

4. Keep the prefix stable so caching works

Providers offer prompt caching, with discounts and limits that vary by model.

Keep reusable content first and changing material later where your provider recommends it. Anthropic's pricing documentation lists cache read and write charges.

Check actual cache usage rather than assuming every repeated prefix hits.

5. End sessions, do not extend them

Long sessions can accumulate earlier messages and irrelevant file reads. When the task changes, consider a new session with a short summary of the state you still need.

Some tools offer compaction, but summarizing can lose details and saving depends on the context and caching behavior.

Compare usage before and after rather than applying a fixed compaction threshold.

6. Send less data

Try minifying JSON with our formatter before sending it.

Count the exact versions with your model tokenizer. Trim irrelevant log lines and avoid unnecessary encoded data, while preserving the facts the task needs.

Compare languages on your actual text rather than assuming a fixed ratio.

What it adds up to: a worked estimate

This is not a measured benchmark. Real usage varies too much for one number to be honest.

It is a model of a plausible day with the assumptions written down, so you can swap in your own.

Assume a developer runs four agent sessions a day, each around 30 turns. Instruction file, history growth, file reads and reply length are the four things the techniques above change.

Per session:

Component | Untuned | Tuned | What changed

Instruction file, re-sent 30 times | 1,300 × 30 = 39,000 | 400 × 30 = 12,000 | Trimmed to a screen

File reads accumulated in context | ~60,000 | ~20,000 | Grep first, partial reads, no generated files

History re-read across the session | ~250,000 | ~120,000 | Smaller reads and replies compound, compact at midpoint

Replies (output) | ~15,000 | ~8,000 | No wrapping, diffs not whole files

Total per session | ~364,000 | ~160,000 | About 56 percent less

Then apply caching. In the untuned case assume little of the prefix is stable and 30 percent of input is served from cache.

In the tuned case the instruction file and early context are stable and 70 percent is cached.

Charging cached input at a tenth of the normal rate, the two assumed cache shares produce the effective input multipliers shown below. Actual cache writes and hits can change the result.

Put together across four sessions a day and about 230 working days a year:

| Untuned | Tuned

Raw tokens per year | ~335 million | ~147 million

Effective billed input after caching | ~73 percent of raw | ~37 percent of raw

Effective cost, relative, output priced at five times input | 100 | ~29

Under these assumptions the tuned setup costs a bit under a third of the untuned one for the same work.

Change the assumptions and the number moves, but not the shape: the savings come from re-reading less and caching more, and they compound because repeated context can contribute to later calls.

If you are on a subscription with a usage cap rather than paying per token, the subscription may count usage differently, so verify its rules before applying this estimate.

What not to do

Do not starve the model of context it needs. An agent that cannot see the relevant code guesses, and a wrong guess costs more turns than the read would have. The goal is no waste, not no reading.

Do not turn off reasoning to save tokens on hard problems. Thinking tokens are expensive and they are the reason the answer is right. Save them on easy tasks, spend them on hard ones.

Do not compress your instructions into cryptic shorthand. A model reads terse, clear English fine. It reads your private abbreviations badly and does more turns to compensate.

Check the input and output split. Optimize the part that costs more in your own sessions.

The five line version

Keep necessary instructions clear, search before reading whole files, ask for a useful reply format and measure caching. Compare these changes on your own sessions before claiming a saving.

For the unit behind the usage report, read how LLM token accounting works.