What one run costs
A typical run of this task sends about 60K tokens and gets back 1K, over a single call — 61K tokens in total. Those are the numbers below, against the published prices.
| Model | Vendor | Per run | Can do what this needs? | Price evidence |
|---|---|---|---|---|
| GPT-5 Nano | OpenAI | $0.00340 | not stated | Vendor page |
| Gemini 2.5 Flash Lite | $0.00640 | not stated | Vendor page | |
| GPT-4.1 nano | OpenAI | $0.00640 | confirmed | Corroborated |
| Claude Haiku 5.5 | Anthropic | $0.00650 | confirmed | Vendor page |
| GPT-6 Luna | OpenAI | $0.00650 | not stated | Vendor page |
| GPT-4o mini | OpenAI | $0.00960 | confirmed | Corroborated |
| Jamba Mini | AI21 | $0.0124 | confirmed | Vendor page |
| GPT-5.6 Luna | OpenAI | $0.0132 | not stated | Vendor page |
Lots in, little out, and no chain of reasoning to get wrong — which makes this the task where a cheap model with a big context window (The total tokens a model can consider at once: your prompt, the conversation so far, and the answer it is writing.) is not a compromise.
Two things improve the result more than a bigger model. Say what the summary is for, since a summary for a decision and a summary for a file note are different documents. And ask for the parts it is unsure about, which surfaces what the document did not actually say.
Beware the middle. Models attend less to material buried halfway through a long context, so for anything load-bearing, ask specifically about it rather than trusting it to surface.
What we would pick
This section is our judgement, not a figure read off a page. Everything above is arithmetic on published prices; this is an opinion, and it is labelled as one.