Running an agent without a surprise bill
- The context grows on every pass, so the last turn costs several times the first.
- Each growth invalidates the cache, which is why writes dominate agentic bills.
- Give it a dry run before you give it authority.
- Cap the loop. An agent that cannot stop is a bill that cannot stop.
Why the twentieth turn is not like the first
An agent is a model in a loop: decide, call a tool, read the result, decide again. The result of each step is appended to the context, so the context the model reads grows on every pass.
Turn one might send five thousand tokens. Turn twenty might send eighty thousand, because it is carrying nineteen tool results. You are charged for all of it, every time. The cost of a loop is not linear in its length — it is closer to quadratic.
And that is also why caching behaves badly here
Because the context changes on every pass, the cached prefix is invalidated on every pass. An agent is close to the worst case for a prompt cache (Paying a reduced rate for a prefix the vendor has already processed, instead of full price for sending it again.): it rewrites constantly and reads each entry once.
This is the main reason agentic workloads are dominated by cache-write (The charge for putting a prefix into the prompt cache. Usually more than the input rate, not less.) charges rather than by output. If your tool lets you keep the stable part of the context — the system prompt (Standing instructions sent ahead of every conversation, setting the role, rules and format.), the files that are not changing — separate from the growing part, that is the single most effective thing you can do about the bill.
Dry run first
An agent that has misunderstood the goal pursues the wrong thing energetically and without stopping to check. Before an agent gets authority over anything that matters, have it report what it would do and read the list.
This costs one run and catches the class of failure that is otherwise discovered afterwards.
Cap everything
Set a maximum number of turns. Set a spend limit if the platform offers one. Have the loop stop and ask rather than retry indefinitely when a tool keeps failing — a retry loop around a permission error will run until something else stops it.
None of this is pessimism about agents. It is the same discipline as a timeout on a network call: the useful behaviour and the runaway behaviour look identical for the first few seconds.