Learn · 4 of 7
What it costs, and why the surprise
Where the money actually goes, which is rarely where people expect.
Prices are quoted per million tokens (How model prices are quoted: dollars per million tokens, charged separately for what you send and what comes back.), in each direction, and a million tokens is less than it sounds: roughly one mid-sized codebase, or one long afternoon with a coding agent.
The spread between models is about six hundred times from cheapest to dearest, which is the first surprise. The second is where your own money actually goes.
Three things dominate, in this order
Context, re-sent. Nothing persists between calls, so every turn re-sends the whole conversation. A 100,000-token codebase across forty turns is four million input tokens. The answers in the same session might total eighty thousand. The context is fifty times the output.
Cache writes (The charge for putting a prefix into the prompt cache. Usually more than the input rate, not less.). If you are caching that context — and you should be — writing an entry costs more than sending it normally: around twice the input rate for a one-hour entry. Rewrite it every turn and you pay the premium every turn.
Output, last. It has the highest rate per token, but on most real workloads there is far less of it than people assume.
That ordering is why “use a cheaper model” is often the wrong first move, and “send less context” is usually the right one.
The model that costs nothing extra
Training a frontier model (The largest and most capable model a vendor currently offers, priced accordingly.) costs tens or hundreds of millions of dollars. You are never charged for any of it. What you pay for is inference (Running a trained model to get an answer. Everything you pay for by the token is inference, not training.) — running the finished model — and no amount of your usage teaches it anything, because the weights are fixed between releases.
Working it out rather than guessing
The calculator takes a real workload — tokens, calls a day, cache behaviour — and prices it against every model at once. Where a vendor does not publish a figure it needs, it refuses to produce a number rather than filling the gap, which is the whole point.