You pay for the tokens you send and the tokens a model returns. Trim a prompt only when the remaining text still contains the facts needed to answer it.

These six changes help you find that boundary; measure token use and answer quality separately.

1. Trim the question, not the information

Here is a bug report written the way people write to a colleague, with a greeting, an apology, a thank you and the question repeated twice. Then the same report with only the code and the question.

Original question:
Hi, could you help? Why does JSON.stringify drop the email field here?
const user = { name: "Ada", email: undefined };
Thanks for looking at this. Why is email missing?

Shorter question:
Why does JSON.stringify omit email from this object?
const user = { name: "Ada", email: undefined };

Both questions include the same code and the missing field.

The shorter question drops the greeting and repetition. Count each exact string with your tokenizer, then verify that the response still explains the undefined value rather than a different issue.

2. Send the data in its cheapest shape

For flat records, compare pretty printed JSON, minified JSON and CSV.

JSON retains field names and data types; CSV needs a header and a convention for values such as booleans, nulls and strings containing commas.

JSON:
[{"id":1,"name":"Ada","active":true}]

CSV:
id,name,active
1,Ada,true

Minifying removes layout whitespace, while CSV puts each field name in the header.

Neither gives a universal token reduction or guarantees equal accuracy. Our JSON formatter lets you compare a compact JSON version.

Keep nested data in JSON unless you have a clear rule for flattening it without losing information.

3. Send the slice, not the file

Start with the error lines from a log, then include the surrounding request IDs, timestamps and events needed to explain them. Filtering everything else is useful only if it preserves the evidence for the question.

grep -E "ERROR|WARN" application.log

Run that filter as a first pass, then inspect nearby lines before sending the result.

Anthropic reports a separate data task in its cost guide: 25 aggregate questions cost $5.01 with a CSV pasted into the prompt and answered 6 correctly. Uploading the file and querying it with code cost $0.40 and answered all 25.

Those results describe that test, not a guarantee for your logs.

4. Drop the boilerplate

Review repeated system instructions and examples. Remove greetings and redundant rules, while keeping constraints your application depends on.

Anthropic's prompt engineering guide gives ways to make the task explicit; test those changes on your model.

Examples can help define an unusual format or clarify a difficult label. Compare a prompt with and without them on representative inputs.

Remove an example only after checking the cases it was meant to explain, including inputs that are easy to misclassify.

If you use an agent with an instruction file, the same logic applies to the file, and we wrote a whole post on keeping instruction files short.

5. Ask for the shape of the answer

For the following models, output tokens cost more than input. On Anthropic's price sheet Claude Sonnet 5.5 is $2 per million tokens in and $10 out, and Claude Opus 5.5 is $4 in and $20 out, five times the input price in both cases.

So a model that answers a yes or no question with three paragraphs is spending the expensive kind of token on words you scroll past.

The fix is to say what shape you want. "Reply with one word." "Give the fix, no explanation." "Return only the JSON." Anthropic's cost documentation makes the same point: to shorten visible responses, specify the exact output shape in the prompt, ideally with an example.

It also warns against the tempting shortcut, a low max_tokens cap.

The model never sees that number, so it does not write shorter, it gets cut off mid sentence and you pay for a retry.

For reasoning tasks there is a measured version of this.

In February 2025 a team at Zoom published Chain of Draft, which replaces "think step by step" with one extra sentence: think step by step, but keep a minimum draft for each step, with five words at most. Their numbers on grade school maths, from the paper's own table:

GSM8K                    accuracy   avg tokens
GPT-4o, chain of thought    95.4%       205.1
GPT-4o, chain of draft      91.1%        43.9
Claude 3.5 Sonnet, CoT      95.8%       190.0
Claude 3.5 Sonnet, CoD      91.4%        39.8

Read that honestly.

On arithmetic the short drafts gave up about four points of accuracy for roughly 80 percent fewer tokens. On their sports understanding task the same trick on Claude 3.5 Sonnet went from 189.4 tokens at 93.2 percent to 14.3 tokens at 97.3 percent, so fewer tokens and a better score.

Whether the trade is worth it depends on your task, which is why you measure rather than assume.

The lesson that holds everywhere is that the length of the visible reasoning is a knob you control, and the useful setting depends on the task.

6. Cache the repeated prefix

Repeated system instructions, tool definitions and documents are candidates for prompt caching.

A cache hit can reduce the price of input already processed. It does not reduce the number of tokens you send or guarantee a hit for every request.

Anthropic. A cache hit is billed at a tenth of the input price on most models, at a twentieth on Claude Opus 5.5, and writing the cache costs 1.25 times the input price for a five minute cache. The arithmetic on the pricing page is that caching pays for itself after one hit. Their cost guide measured agent loops running 2.7 to 5.3 times cheaper with caching, with real traffic reading a median of 84 percent of input tokens from cache.

OpenAI. Caching is automatic for prompts of 1024 tokens or more, matched on the exact repeated prefix in steps of 128 tokens. The cached price is 50 percent off on GPT-4o and up to 90 percent off on newer models, per OpenAI's caching guide.

Google. Gemini 2.5 and newer cache implicitly once a prompt passes the model's minimum, 2048 tokens on 2.5 Flash and Pro. On the price sheet cached input on 2.5 Flash is $0.03 per million against $0.30 standard, a 90 percent cut, with a separate hourly storage charge for explicit caches.

Keep reusable content at the beginning where the provider recommends it, and put changing questions later.

Cache mechanisms and minimum sizes differ. Check the usage fields in the response to see how many tokens were actually served from cache.

Use batch pricing when the task can wait

Batching.

If nobody is waiting for the answer, Anthropic and Google both sell a batch mode at 50 percent off every token, and on Anthropic the batch discount stacks on top of the cache discount.

It is not a way to send fewer tokens, it is a way to pay less for the same ones, but for nightly jobs, evaluations and backfills it is a pricing option to compare for work that can wait.

Putting it together

Start with one recurring request. Keep the original input and its checked answer, then change one thing: prompt wording, data shape, selected lines or reply format.

Compare token counts and correctness across the same test cases.

If the answer loses needed evidence, restore that context.

Paste your JSON into our formatter, compare its token count before and after minifying, and verify the model response before keeping the smaller input.