Keeping the bill down
- Output usually costs three to five times input — long answers cost more than long questions.
- Try the smallest model in a family first and move up only when it visibly fails.
- Reasoning models bill you for their working-out, whether or not you see it.
- A million tokens is one afternoon with a coding agent, not a year's supply.
Output is the expensive half
Prices are quoted in both directions and they are not the same number. A model at $3 in and $15 out charges five times more for what it writes than for what it reads.
The practical consequence is that asking for brevity saves real money, and asking a model to “explain your reasoning in detail” on a task that does not need it is a way of paying five times over for words nobody reads. Cap the output length where the task has a natural one.
Start small, move up on evidence
Every vendor ships a range, and the flagship is typically five to twenty times the price of the smallest model in the same family. On a great deal of real work — classification, extraction, rewriting, summarising — the small model is indistinguishable.
The habit worth building is to try the cheapest model in the family first and move up only when it measurably fails at your task, rather than defaulting to the one in the headlines. Measurably is the operative word: keep twenty examples you care about and compare.
Reasoning models bill for thinking
A reasoning model (A model trained to produce a long internal working-out before its answer. More accurate on hard problems, slower and dearer.) generates a long internal working-out before its answer, and those intermediate tokens are output tokens charged at the output rate — even though most interfaces never show them to you.
That can be five or ten times the cost of a conventional model on the same question, and it is worth it when being wrong is expensive. For drafting an email it is an expensive way to be slower.
Where the tokens actually go
People consistently underestimate how much context costs and overestimate how much the answer costs. A 100,000-token codebase re-sent across forty turns is four million input tokens; the answers across the same session might be eighty thousand. The context is fifty times the output.
That is why caching matters more than brevity on anything agentic, and why pruning the context beats shortening the reply.
This is approximate and cannot be otherwise. Every model carves text up with its own fixed table, so the same sentence is a different number of tokens to each one. This uses about four characters per token, which is a decent rule of thumb for ordinary English prose and a poor one for code, for rare words, and for languages the tokenizer saw little of. For a bill, the only authority is the vendor's own count.