Pull structured data out of documents
“I need the numbers out of a pile of invoices or forms”
What one run costs
A typical run of this task sends about 15K tokens and gets back 800, over a single call — 15.8K tokens in total. Those are the numbers below, against the published prices.
| Model | Vendor | Per run | Can do what this needs? | Price evidence |
|---|---|---|---|---|
| GPT-5 Nano | OpenAI | $0.00107 | not stated | Vendor page |
| Gemini 2.5 Flash Lite | $0.00182 | not stated | Vendor page | |
| Claude Haiku 5.5 | Anthropic | $0.00190 | confirmed | Vendor page |
| GPT-6 Luna | OpenAI | $0.00190 | not stated | Vendor page |
| GPT-5.6 Luna | OpenAI | $0.00396 | not stated | Vendor page |
| GPT-5.4 Nano | OpenAI | $0.00400 | not stated | Vendor page |
| Gemini 3.1 Flash Lite | $0.00495 | not stated | Vendor page | |
| GPT-5 Mini | OpenAI | $0.00535 | not stated | Vendor page |
A volume task, which makes the price per document the whole decision — and at this tier a thousand documents costs a few dollars.
What actually determines the quality is the prompt, not the model. Give one worked example of the exact output shape, ask for strict JSON, and set the temperature (How much randomness to allow when picking each token. Low is repeatable and dull; high is varied and unreliable.) as low as the tool allows so the same document gives the same answer twice.
Build a way to spot failures before you run a thousand. A field that is silently empty on three per cent of documents is much worse than one that errors.
What we would pick
This section is our judgement, not a figure read off a page. Everything above is arithmetic on published prices; this is an opinion, and it is labelled as one.