Working inside a context window
- A context window is a budget spent on every call, not a memory that accumulates.
- The whole conversation is re-sent each turn — that is why long chats get expensive.
- Models attend least to whatever is buried in the middle of a long context.
- Hitting the limit does not error; most tools silently drop your oldest turns.
The thing almost everyone gets wrong
A context window (The total tokens a model can consider at once: your prompt, the conversation so far, and the answer it is writing.) sounds like memory. It is a budget.
Nothing persists between calls. Each time you send a message, your tool re-sends the system prompt (Standing instructions sent ahead of every conversation, setting the role, rules and format.), every file you attached, and the entire conversation so far, as one block of text. The model reads all of it from scratch, answers, and forgets it again. The next turn re-sends everything plus the answer it just gave.
That single fact explains most of what feels strange about working with these tools (Letting a model call functions you define — search, read a file, send a request — and read back what they return.). Long chats get slower and more expensive with every turn, not because the model is tiring but because you are paying to re-read the transcript each time. An assistant “forgets” a decision from twenty messages ago because that part of the transcript was quietly dropped to make room. And a prompt cache (Paying a reduced rate for a prefix the vendor has already processed, instead of full price for sending it again.) exists at all because re-sending the same prefix every turn is so obviously wasteful.
Four things come out of one allowance
The system prompt, your attached files, the conversation so far, and the space the answer needs. Spend too much on the first three and there is no room for the fourth.
Set the slider for attached files high enough and watch the bar: what happens at the limit is not an error. Most tools trim the oldest turns and carry on, and nothing tells you they did.
Where to put what matters
Models use the beginning and the end of their context well and are measurably worse at the middle. This is consistent enough to plan around.
Put reference material the model must not miss at the start. Put the instruction you actually want followed at the end — last, after the documents, not before them. If something in the middle is load-bearing, ask about it directly rather than trusting it to surface on its own.
Pruning beats a bigger window
A larger context window is not the same as a model that uses one well, and the habit worth building is removing things rather than adding them. Start a fresh conversation when the subject changes. Paste the three relevant functions instead of the whole file. Summarise a long thread into a paragraph and carry the paragraph forward.
It is cheaper, and it measurably improves the answers.