Every model bill is the same two numbers: tokens you send and tokens you get back. A lot of what gets sent is not needed to produce the answer, and a lot of what comes back is not needed to use it. This post is a list of the places where the same question can be asked with fewer tokens and still return the same answer. Every number here is either measured with a real tokenizer or quoted from a paper or a vendor price sheet, with the source linked. If tokens are new to you, start with how models read text, then come back.
The counts below were made with OpenAI's open source tokenizer, the o200k encoding used by the GPT-4o and GPT-5 families. Claude and Gemini split text differently, so the exact numbers differ there, but the ratios hold because the savings come from sending less text, not from how the text is split.
1. Trim the question, not the information
Here is a bug report written the way people write to a colleague, with a greeting, an apology, a thank you and the question repeated twice. Then the same report with only the code and the question.
Verbose version, with greeting and thanks: 148 tokens
Code plus one line question: 28 tokens
Saving: 81%
Nothing the model needs was removed. The function is there, the symptom is there, the question is there. The 120 tokens that went away were politeness, and the model does not answer better for it. This is the cheapest habit to change because it costs nothing and it applies to every single message.
2. Send the data in its cheapest shape
The same 50 records, each with an id, a name, an email, a flag and a score, cost very different amounts depending on how they are written.
Pretty printed JSON, two space indent: 2202 tokens
Minified JSON, no whitespace: 1302 tokens
CSV with a header row: 657 tokens
Minifying alone saves 41 percent, because every indent and newline is a token. CSV saves 70 percent, because the keys are written once in the header instead of once per record. The model reads all three equally well for flat records. Our JSON formatter minifies with one click, and the tokens post explains why JSON punctuation is so expensive in the first place. The exception is deeply nested data, where CSV stops being honest and minified JSON is the right stop.
3. Send the slice, not the file
An 801 line log file with one error in the middle, and the same file filtered to the lines that contain ERROR or WARN before it is sent.
Whole log file: 24838 tokens
Only the ERROR and WARN lines: 38 tokens
That is a 650 to 1 difference for the same diagnosis, because the 800 INFO lines carried nothing the question needed. The general rule: if you can find the relevant part with grep, a search, or a database query, do that first and send the result. Anthropic's own cost guide reports a sharper version of the same idea for tabular questions. In their test, 25 aggregate SQL style questions over a data file cost $5.01 with the file pasted into the prompt and got 6 right, and cost $0.40 with the file uploaded and queried by code and got all 25 right. Pasting was 12 times more expensive and worse.
4. Drop the boilerplate
Two places where people pay every request for words that do nothing. First, the system prompt. A typical block of reassurance, be helpful, be accurate, think carefully, be concise but thorough, is 92 tokens. The three rules that actually change behaviour, answer directly, say when unsure, plain text unless asked, are 17. Modern models already do the rest by default, and Anthropic's prompt guidance for its current models says that over prescriptive prompts written for older models tend to reduce output quality, not raise it.
Second, examples. A sentiment classifier with nine worked examples is 179 tokens. The same request with zero examples and the rule "reply with one word" is 28. For a task this simple the examples buy nothing. For a hard or unusual output format they still earn their place, so the honest rule is: start with none, add one only when the output is wrong without it, and never add a third because two felt thin.
If you use an agent with an instruction file, the same logic applies to the file, and we wrote a whole post on keeping instruction files short.
5. Ask for the shape of the answer
Output tokens cost more than input. On Anthropic's price sheet Claude Sonnet 5.5 is $2 per million tokens in and $10 out, and Claude Opus 5.5 is $4 in and $20 out, five times the input price in both cases. So a model that answers a yes or no question with three paragraphs is spending the expensive kind of token on words you scroll past.
The fix is to say what shape you want. "Reply with one word." "Give the fix, no explanation." "Return only the JSON." Anthropic's cost documentation makes the same point: to shorten visible responses, specify the exact output shape in the prompt, ideally with an example. It also warns against the tempting shortcut, a low max_tokens cap. The model never sees that number, so it does not write shorter, it gets cut off mid sentence and you pay for a retry.
For reasoning tasks there is a measured version of this. In February 2025 a team at Zoom published Chain of Draft, which replaces "think step by step" with one extra sentence: think step by step, but keep a minimum draft for each step, with five words at most. Their numbers on grade school maths, from the paper's own table:
GSM8K accuracy avg tokens
GPT-4o, chain of thought 95.4% 205.1
GPT-4o, chain of draft 91.1% 43.9
Claude 3.5 Sonnet, CoT 95.8% 190.0
Claude 3.5 Sonnet, CoD 91.4% 39.8
Read that honestly. On arithmetic the short drafts gave up about four points of accuracy for roughly 80 percent fewer tokens. On their sports understanding task the same trick on Claude 3.5 Sonnet went from 189.4 tokens at 93.2 percent to 14.3 tokens at 97.3 percent, so fewer tokens and a better score. Whether the trade is worth it depends on your task, which is why you measure rather than assume. The lesson that holds everywhere is that the length of the visible reasoning is a knob you control, and the default setting is long.
6. Reuse the prefix instead of resending it
Most requests to a model start the same way: the same system prompt, the same tool definitions, the same document, the same conversation so far. Every vendor now charges less for a prefix it has already seen, and this is the biggest single saving on the list, because it does not change what you send, only what you pay.
- Anthropic. A cache hit is billed at a tenth of the input price on most models, at a twentieth on Claude Opus 5.5, and writing the cache costs 1.25 times the input price for a five minute cache. The arithmetic on the pricing page is that caching pays for itself after one hit. Their cost guide measured agent loops running 2.7 to 5.3 times cheaper with caching, with real traffic reading a median of 84 percent of input tokens from cache.
- OpenAI. Caching is automatic for prompts of 1024 tokens or more, matched on the exact repeated prefix in steps of 128 tokens. The cached price is 50 percent off on GPT-4o and up to 90 percent off on newer models, per OpenAI's caching guide.
- Google. Gemini 2.5 and newer cache implicitly once a prompt passes the model's minimum, 2048 tokens on 2.5 Flash and Pro. On the price sheet cached input on 2.5 Flash is $0.03 per million against $0.30 standard, a 90 percent cut, with a separate hourly storage charge for explicit caches.
All three work the same way under the hood: the cache matches the start of the prompt byte for byte. One timestamp, one request id, one "today is" at the top of your system prompt and nothing after it ever hits the cache. Put the stable text first and the changing text last, and check the usage field in the response to confirm the cached token count is not zero.
The one that is not on the list
Batching. If nobody is waiting for the answer, Anthropic and Google both sell a batch mode at 50 percent off every token, and on Anthropic the batch discount stacks on top of the cache discount. It is not a way to send fewer tokens, it is a way to pay less for the same ones, but for nightly jobs, evaluations and backfills it is the easiest half off you will ever get.
Putting it together
Take the log question from section 3 as it would normally be sent: the whole file plus a chatty question, about 25,000 tokens in. Filter the file first and trim the question and it is under 70 tokens in. Ask for the fix without a lecture and the answer comes back in a few dozen tokens instead of a few hundred. Run it through a cached system prompt and the little that is left is billed at a tenth. None of those steps changed the question, and none of them changed the answer. They changed what you paid for it.
Measure before you believe any of this on your own prompts. Paste one into our JSON formatter and minify it to see the first saving instantly, and count the rest with the tokenizer your vendor publishes. The numbers in this post came from exactly that.