Pick ChatGPT, Claude, Gemini or Ollama by What You Pay Each Month

By Sheng Pang · Published · 10 min read

AI in practice ai pricing chatgpt claude gemini ollama

Twenty questions a day to GPT-5.6 Sol costs $2.51 a month at API prices. ChatGPT Plus costs $20 for the same model.

We wanted to see where the rest of the $20 goes, so we set up three example months, a chatter, a reader and a script, and priced each one every way the vendors let you pay.

Two ways to buy the same model, and a third that swaps the model

OpenAI, Anthropic and Google each sell a chat app on a flat monthly fee, and each sells the same models per token through an API.

The third way is different. You download an open weight model, which is not any of their models, and run it on your own laptop with Ollama.

Nobody bills you for that, but you need the memory to hold the model and you wait on your own hardware.

ChatGPT Plus is $20, Claude Pro is $20, Google AI Pro is $19.99.

OpenAI and Google also sell a cheaper tier, ChatGPT Go at $8 and Google AI Plus at $4.99. Anthropic doesn't, its next step up from Free is Pro, and all three sell a $100 tier with higher limits.

One question is 35 tokens and one good answer is 202

A model bills you per token, and in English a token is a bit less than a word you type.

We didn't want to trust the rule of thumb, so we ran some of our own articles through OpenAI's o200k tokenizer. Our 1,079 word Ollama guide came out at 1,365 tokens, a 1,308 word one at 1,573 and a 1,446 word one at 1,711.

Call it 1.2 tokens per word, the ratio we use below, so you can redo any sum here with your own word count.

We counted with OpenAI's tokenizer, and Anthropic's pricing page says Claude models from 4.7 on produce about 30 percent more tokens for the same text.

The test question was this one, and you can paste it into any of the apps yourself.

Why does my Python script say list index out of range when the list has 3 items and I use index 3? Here is the line: print(items[3])

The question is 35 tokens.

We wrote a 157 word answer by hand, the kind of reply you would be happy with, and it is 202 tokens. Our token explainer covers why answers cost more than questions.

Twenty questions a day

Twenty questions like the Python one every day is 600 a month, 21,000 tokens in and 121,200 out.

GPT-5.6 Luna is the model free ChatGPT users get, and at OpenAI's API price of $0.20 per million tokens in and $1.20 out the month costs 15 cents. GPT-5.6 Sol, the Plus default, is $4 in and $20 out and costs $2.51 for the same month.

Claude Opus 5.5 has the same two prices as Sol, so $2.51 again, and the Sonnet, Haiku and Gemini rows are in the table further down.

To spend $20 on Sol this way you would have to ask 4,785 questions in the month. That is 160 a day.

The sum assumes every question you ask stands on its own. A chat app resends the whole conversation on every turn, so a long back and forth multiplies the input side.

Long documents cost less than short questions

The second example is pasting a long document like a lease and asking for a summary.

We set it at twice a day and 5,000 words each, about ten pages, so the input side is big and the output stays small.

At 1.2 tokens per word a document is 6,000 tokens in, and a summary is about 400 out, so sixty of them in a month is 360,000 in and 24,000 out.

This costs less than the questions on every model even though each call sends 170 times more input, because input is cheaper than output and the summaries are between a quarter and just over a third of the bill.

Sol and Opus 5.5 cost $1.92 for the month, Haiku 4.5 costs 48 cents and Gemini 3.5 Flash-Lite costs 17 cents.

The script that runs all day

Suppose you wrote a script that tags every note and receipt you save with a model call, 200 calls a day at 500 tokens in and 50 out each.

That comes to 6,000 calls a month, 3 million tokens in and 300,000 out, and here the input side is most of the bill.

Model, API price per million in and outQuestionsDocumentsScript
GPT-5.6 Luna, $0.20 and $1.20$0.15$0.10$0.96
Gemini 3.5 Flash-Lite, $0.30 and $2.50$0.31$0.17$1.65
Gemini 3.8 Flash, $0.75 and $3.75$0.47$0.36$3.38
Claude Haiku 4.5, $1 and $5$0.63$0.48$4.50
Claude Sonnet 5.5, $2 and $10$1.25$0.96$9.00
Gemini 3.1 Pro, $2 and $12$1.50$1.01$9.60
GPT-5.6 Sol, $4 and $20$2.51$1.92$18.00
Claude Opus 5.5, $4 and $20$2.51$1.92$18.00

Sol and Opus 5.5 cost $18 here, nearly what you pay for Plus.

But the subscription is no use to a script because the chat apps have no endpoint you can call.

Luna costs 96 cents for the same month and Flash-Lite costs $1.65.

We didn't run this script, so whether the cheap tags are good enough is something you would have to check by reading them.

Google lists Gemini 3.8 Flash at $0.75 and $3.75 through December 31, 2026 and $1.50 and $7.50 from January 1, 2027, so that row doubles next year. Gemini bills thinking tokens as output on top.

On a laptop the bill is electricity

We used Qwen3 8B because it was already on the laptop, a 2023 MacBook Pro with an M2 Pro chip and 32 GB of memory.

The first run produced nothing. Qwen3 thinks before it answers, we had capped the output at 400 tokens and the thinking used all 400.

With thinking turned off it read the prompt in 0.2 seconds and wrote a correct 266 token answer in 10.9 seconds. About 24 tokens a second, and 11.2 seconds for the whole call once you count Ollama's own overhead.

Ollama reported the prompt as 51 tokens, not 35, because Qwen3 has its own tokenizer and wraps the question in its chat template.

We only tested the Python question, so this says nothing about how it would handle a 5,000 word lease, and we didn't try.

Apple ships that laptop with a 140 watt adapter, so it can't draw more than that whatever you run.

600 questions at 11.2 seconds each is 1.9 hours, or 0.26 kWh at the full 140 watts.

The US average home rate was 18.31 cents per kWh in July 2026. So the month of questions costs 5 cents at most.

For the script we scaled from the measured 255 tokens a second for reading and 24 for writing, which makes a 500 in and 50 out call about 4 seconds, 6.7 hours for the month and 17 cents at most.

The real draw is lower than 140 watts, measuring it needs root access we don't have.

And we only timed one model on one laptop, the first run with thinking on reported 31 tokens a second and the second without reported 24 and we don't know why the second was slower.

The download was 5.2 GB, and you need the memory to hold it. Ollama's model pages say at least 8 GB of RAM for a 7B class model, 16 GB for 13B and 64 GB for 70B.

The $20 plans do not publish a message count

None of the three vendors says how many messages you get for the $20.

OpenAI says your limit depends on the model and the task. Anthropic says Pro gets more usage per five hour session than Free and Max gets 5 or 20 times Pro. Google says AI Pro gets four times the free limits.

What they do list is features, file uploads, voice, image generation, memory across chats and on Google's side 5 TB of storage.

Free got better in August 2026 as well, when OpenAI removed the text message limit for free users and made GPT-5.6 Luna the default. Files and voice still have limits.

How we would choose

Within the API, look at the output price first.

An output token costs five times an input token on most of these models, and an answer has more tokens than the question. In the questions example the answers cost 29 times what the questions did on those models, and even more on Luna and Flash-Lite.

A few chats a day and no code means a free plan is probably enough. That is the Luna row of the table, and the free ChatGPT app gives you that same model.

What the table can't tell you is how much better Sol or Opus answer than Luna or Haiku. We didn't measure quality here, and if the cheap model's answers annoy you, the $20 buys the better model inside the app with no setup.

The messy case is the person who chats a lot and pastes a file a couple of times a week. We priced 600 questions plus ten documents a month and it's $2.83 on Sol, so the API still wins on paper. The upload box in the app is worth something though and we have no way to price it.

So pay the $20 when the free app's upload and message limits get in your way, or when you want the better model without setting anything up, and pay it to whichever vendor's app you already open most.

Ollama makes sense when the data can't leave the machine. Nothing we priced there came to more than 17 cents a month, and you wait about 11 seconds per answer instead.

Scripts go on the API. We would start with the smallest model from whichever vendor you already have an account with, since the table says it costs a tenth and reading its output is the only way to learn whether that is enough for you.

Set a spending limit before the first call you make. A loop with a bug keeps calling the model until something stops it, and you pay for every call.

The table hides one cost, the setup you do yourself and the app does for you.

That means an API key, code to call the model and code to handle failures.

Do the same sum for your own month

To price your own month, count one real question and one real answer with a tokenizer, multiply by how many you ask a day and by 30, and put the two totals against the price sheet you plan to use.

If your month lands around the $2 to $3 that the questions and documents columns show, the $20 is paying for the features rather than the model.

We'd also try the laptop. A 16 GB machine runs an 8B model with room to spare by Ollama's guidance, and you can ask it the same question and compare.

Our token saving guide shows how to shrink the input side when the script bill is the problem. Tell us on the Contact page if your month looks nothing like these three.

← More AI in practice articles  ·  All articles

↑ Top