Skip to content

Decision instrument

Which model?

The best model isn't the smartest one.It's the one that fits the job.

28 models from the OpenRouter catalog on 5 Oct 2026. Prices and Artificial Analysis indices are published there. The score beside the recommendation is arithmetic on those figures and on the job you set. The volume is a hypothesis.

Pull structured fields out of a pile of documents.

agentic index

028.957.9$100$1kGLM 5.3 FlashClaude Fable 5.1

Workload cost, log scale

  • Fits the job
  • Leads the agentic index

Claude Opus 5.5, Claude Sonnet 5.5, DeepSeek V4.1 Flash, GPT-6 Luna, GPT-6.1 Sol, Grok 4.7, MiMo-V2.6-Flash, and MiMo-V2.6-Pro sit this job out: no agentic index in this snapshot.

Recommended

GLM 5.3 Flash

Fit 92.5calculated

Fit is this page's score for the job, from the weights below. It is calculated here. It is not a benchmark.

Index
50.9
This workload
$16.00
Context
1.05M

GLM 5.3 Flash fits this document extraction job best of the 28 models on this page, at $16.00.

Claude Fable 5.1 leads the agentic index at 57.9. On this workload it costs $1,200, against $16.00 for GLM 5.3 Flash.

GLM 5.3 is next, at fit 82.6.

Weights
45%
40%
15%

Fit = 45% quality + 40% cost + 15% context

Context adds the same amount to every model that can take this job: each window is at least twice the prompt.

Candidates

Sorted by fit. The index and the list price are published. The workload cost and the fit are calculated.

GLM 5.3 Flash

Z.ai · 1.05M context

Fit
92.5
agentic index
50.9
This workload
$16.00
List price
$0.150 / $0.500 per 1M

GLM 5.3

Z.ai · 1.05M context

Fit
82.6
agentic index
53.1
This workload
$60.00
List price
$0.050 / $7.00 per 1M

GPT-5.6 Luna

OpenAI · 1.05M context

Fit
81.5
agentic index
42.1
This workload
$25.60
List price
$0.200 / $1.20 per 1M

Qwen3.8 27B

Qwen · 1M context

Fit
77.8
agentic index
45.8
This workload
$54.40
List price
$0.425 / $2.55 per 1M

DeepSeek V4 Flash 0423

DeepSeek · 1.05M context

Fit
75.4
agentic index
26.3
This workload
$12.64
List price
$0.030 / $1.28 per 1M

What 10,000 documents would cost

Uncached prompt plus completion, at the catalog rate for this prompt length. The count and the token sizes are a hypothesis. Changing them rescores the page.

  1. DeepSeek V4 Flash 0423$12.64
  2. GLM 5.3 Flash$16.00
  3. DeepSeek V4 Pro 0423$20.04
  4. GPT-5.6 Luna$25.60
  5. MiniMax M3$33.60
  6. Qwen3.8 27B$54.40
  7. Nemotron 3 Ultra$57.60
  8. GLM 5.3$60.00
  9. DeepSeek V4 Pro 0813$72.00
  10. Gemini 3.8 Flash$90.00
  11. Muse Spark 1.2$134
  12. Kimi K3$166
  13. Mistral Medium 3.5$180
  14. Qwen3.8 Max (0902)$208
  15. Grok 4.6$208
  16. GPT-5.6 Sol$240
  17. GPT-5.6 Terra$256
  18. Claude Opus 5$600
  19. Claude Fable 5.1$1,200
  20. GPT-6 Astra$1,200

How the number is made

  1. OpenRouter catalog
  2. Monthly catalog fetch, 5 Oct 2026
  3. Benchmark and price
  4. Workload score
  5. This page
published
Artificial Analysis indices as published on the OpenRouter model catalog. Fetched 5 Oct 2026 from https://openrouter.ai/api/v1/models. Uncached prompt and completion, USD per token. When the catalog publishes a long-prompt override and the input length meets its minimum, that rate is used. Cache reads, batch rates, and web search are not in the cost.
calculated
Quality is this job's published index, divided by the highest index among models that can take the job.
Cost is the uncached dollar total, scored on a log scale from the cheapest eligible model (1) to the dearest (0). The dollars themselves are shown unscaled.
Context is headroom over the prompt. A window twice the prompt scores 1.
hypothetical
The job's volume, token sizes, and opening weights. They are not measurements. Moving them is the decision.

What this page leaves out

  • 28 models: the current shipping line from each provider on this page, and the cheaper sibling when the catalog publishes one. Not the whole catalog. Batch and free aliases stay off. The extract is replaced on the first of each month; a model that leaves the catalog stops the refresh rather than being guessed.
  • Latency. This catalog snapshot does not publish a comparable speed, so speed is not a weight.
  • Cache hits, batch rates, and web search. The cost is a cold prompt.
  • A missing index sits the model out of that job. The absence is the catalog's, and it is not filled in.
  • Fit is relative to the models on this page. A different set would move the score.