Vibedia.
All vendors

Show only the vendors you work with. This applies to every page, and is remembered.

Playbooks

Running an agent without a surprise bill

Why the twentieth turn is not like the first

An agent is a model in a loop: decide, call a tool, read the result, decide again. The result of each step is appended to the context, so the context the model reads grows on every pass.

Turn one might send five thousand tokens. Turn twenty might send eighty thousand, because it is carrying nineteen tool results. You are charged for all of it, every time. The cost of a loop is not linear in its length — it is closer to quadratic.

And that is also why caching behaves badly here

Because the context changes on every pass, the cached prefix is invalidated on every pass. An agent is close to the worst case for a prompt cache (Paying a reduced rate for a prefix the vendor has already processed, instead of full price for sending it again.): it rewrites constantly and reads each entry once.

This is the main reason agentic workloads are dominated by cache-write (The charge for putting a prefix into the prompt cache. Usually more than the input rate, not less.) charges rather than by output. If your tool lets you keep the stable part of the context — the system prompt (Standing instructions sent ahead of every conversation, setting the role, rules and format.), the files that are not changing — separate from the growing part, that is the single most effective thing you can do about the bill.

Dry run first

An agent that has misunderstood the goal pursues the wrong thing energetically and without stopping to check. Before an agent gets authority over anything that matters, have it report what it would do and read the list.

This costs one run and catches the class of failure that is otherwise discovered afterwards.

Cap everything

Set a maximum number of turns. Set a spend limit if the platform offers one. Have the loop stop and ask rather than retry indefinitely when a tool keeps failing — a retry loop around a permission error will run until something else stops it.

None of this is pessimism about agents. It is the same discipline as a timeout on a network call: the useful behaviour and the runaway behaviour look identical for the first few seconds.

The loop an agent runs

Four steps, and a context that grows each pass

The agent loopFour boxes in a cycle — decide, call a tool, observe, done? — with the "no" branch returning to the first box. A bar beneath shows the context growing on each pass. The steps are written out in the list below.Decidestep 1Call a toolstep 2Observestep 3Done?step 4not finished — go round again, carrying one more resultcontext after each passand every pass is charged for all of it
  1. Decide. Read everything so far and choose the next action.
  2. Call a tool. Search, read a file, run a command, send a request.
  3. Observe. The result comes back and is appended to the context.
  4. Done?. If not, go round again — now carrying one more result.
ShareOpen LinkedIn