Decision instrument
Which model?
The best model isn't the smartest one.It's the one that fits the job.
28 models from the OpenRouter catalog on 5 Oct 2026. Prices and Artificial Analysis indices are published there. The score beside the recommendation is arithmetic on those figures and on the job you set. The volume is a hypothesis.
Pull structured fields out of a pile of documents.
agentic index
Workload cost, log scale
- Fits the job
- Leads the agentic index
Claude Opus 5.5, Claude Sonnet 5.5, DeepSeek V4.1 Flash, GPT-6 Luna, GPT-6.1 Sol, Grok 4.7, MiMo-V2.6-Flash, and MiMo-V2.6-Pro sit this job out: no agentic index in this snapshot.
Recommended
GLM 5.3 Flash
Fit 92.5calculated
Fit is this page's score for the job, from the weights below. It is calculated here. It is not a benchmark.
- Index
- 50.9
- This workload
- $16.00
- Context
- 1.05M
GLM 5.3 Flash fits this document extraction job best of the 28 models on this page, at $16.00.
Claude Fable 5.1 leads the agentic index at 57.9. On this workload it costs $1,200, against $16.00 for GLM 5.3 Flash.
GLM 5.3 is next, at fit 82.6.
Candidates
Sorted by fit. The index and the list price are published. The workload cost and the fit are calculated.
GLM 5.3 Flash
Z.ai · 1.05M context
- Fit
- 92.5
- agentic index
- 50.9
- This workload
- $16.00
- List price
- $0.150 / $0.500 per 1M
GLM 5.3
Z.ai · 1.05M context
- Fit
- 82.6
- agentic index
- 53.1
- This workload
- $60.00
- List price
- $0.050 / $7.00 per 1M
GPT-5.6 Luna
OpenAI · 1.05M context
- Fit
- 81.5
- agentic index
- 42.1
- This workload
- $25.60
- List price
- $0.200 / $1.20 per 1M
Qwen3.8 27B
Qwen · 1M context
- Fit
- 77.8
- agentic index
- 45.8
- This workload
- $54.40
- List price
- $0.425 / $2.55 per 1M
DeepSeek V4 Flash 0423
DeepSeek · 1.05M context
- Fit
- 75.4
- agentic index
- 26.3
- This workload
- $12.64
- List price
- $0.030 / $1.28 per 1M
What 10,000 documents would cost
Uncached prompt plus completion, at the catalog rate for this prompt length. The count and the token sizes are a hypothesis. Changing them rescores the page.
- DeepSeek V4 Flash 0423$12.64
- GLM 5.3 Flash$16.00
- DeepSeek V4 Pro 0423$20.04
- GPT-5.6 Luna$25.60
- MiniMax M3$33.60
- Qwen3.8 27B$54.40
- Nemotron 3 Ultra$57.60
- GLM 5.3$60.00
- DeepSeek V4 Pro 0813$72.00
- Gemini 3.8 Flash$90.00
- Muse Spark 1.2$134
- Kimi K3$166
- Mistral Medium 3.5$180
- Qwen3.8 Max (0902)$208
- Grok 4.6$208
- GPT-5.6 Sol$240
- GPT-5.6 Terra$256
- Claude Opus 5$600
- Claude Fable 5.1$1,200
- GPT-6 Astra$1,200
How the number is made
- OpenRouter catalog
- Monthly catalog fetch, 5 Oct 2026
- Benchmark and price
- Workload score
- This page
- published
- Artificial Analysis indices as published on the OpenRouter model catalog. Fetched 5 Oct 2026 from https://openrouter.ai/api/v1/models. Uncached prompt and completion, USD per token. When the catalog publishes a long-prompt override and the input length meets its minimum, that rate is used. Cache reads, batch rates, and web search are not in the cost.
- calculated
- Quality is this job's published index, divided by the highest index among models that can take the job.
- Cost is the uncached dollar total, scored on a log scale from the cheapest eligible model (1) to the dearest (0). The dollars themselves are shown unscaled.
- Context is headroom over the prompt. A window twice the prompt scores 1.
- hypothetical
- The job's volume, token sizes, and opening weights. They are not measurements. Moving them is the decision.
What this page leaves out
- 28 models: the current shipping line from each provider on this page, and the cheaper sibling when the catalog publishes one. Not the whole catalog. Batch and free aliases stay off. The extract is replaced on the first of each month; a model that leaves the catalog stops the refresh rather than being guessed.
- Latency. This catalog snapshot does not publish a comparable speed, so speed is not a weight.
- Cache hits, batch rates, and web search. The cost is a cold prompt.
- A missing index sits the model out of that job. The absence is the catalog's, and it is not filled in.
- Fit is relative to the models on this page. A different set would move the score.