Vibedia.
All vendors

Show only the vendors you work with. This applies to every page, and is remembered.

Playbooks

Prompt caching, and when it costs you money

What it actually does

An agent re-sends the same system prompt (Standing instructions sent ahead of every conversation, setting the role, rules and format.) and the same files on every turn. Prompt caching (Paying a reduced rate for a prefix the vendor has already processed, instead of full price for sending it again.) lets the vendor keep the processed form of that prefix for a short while and charge you far less for reading it back — often around a tenth of the input rate.

So far this is straightforwardly good. The part that catches people out is the other half.

Writing is dearer than not caching

Putting something into the cache costs more than simply sending it. Vendors that publish the figure charge roughly 1.25 times the input rate for an entry that lives five minutes, and roughly twice the input rate for one that lives an hour.

That makes caching a bet. Read the prefix several times before it expires and you win comfortably. Rewrite it every turn and you are paying a premium on every single call for a saving you never collect.

It expires on a clock

The window runs from the write, not from the last read. Five minutes means five minutes of wall time, whether you sent ten turns in that period or none.

This is what the timeline above is for. Leave the gap at four minutes with a one-hour cache and almost every turn is a hit. Push the gap past the window and every turn becomes a write — and the total goes above what you would have paid with no cache at all.

The shape that loses money is not exotic. It is a person thinking between turns.

When it is worth it

Caching pays when a large prefix stays identical and is read often and soon: an agent working through a codebase, a support bot with a long standing instruction, a batch job sharing one set of examples across many documents.

It does not pay when the prefix changes on every turn, when the gaps between turns are long, or when the prefix is small enough that the arithmetic is noise either way.

One warning about the figures

Most vendors do not publish a cache-write (The charge for putting a prefix into the prompt cache. Usually more than the input rate, not less.) price at all. Where ours is estimated the widget says so, and the cost calculator refuses to produce a total rather than quietly filling the gap. Fourteen of the models in the catalogue publish a one-hour cache-write price. The other fifty do not.

When the cache pays, and when it does not

52 models publish enough to model this

Cache hits, misses and writes across a conversationOne mark per turn along a time axis. A filled mark is a turn served from cache; an outlined mark is a turn that had to write the prefix again because the cache had expired. The running total is below, and the same figures are in the table.

The same figures as a table
ApproachCache writesCost for the conversation
ShareOpen LinkedIn