Model decision table
Rankings with the tradeoffs left in.
Scan the full catalog, narrow it to your workload, and compare up to three models. Specifications are source-reviewed; fit scores are editorial and their weighting is explicit.
Full catalog
Compare models
Showing 25 models
| Compare | # | Model | Editorial fit | Coding | Reasoning | Tool use | Context | Price / 1M tokens | Evidence |
|---|---|---|---|---|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI · AvailableComplex production workflows · Coding | 9.9editorial | 9.9 | 9.9 | 9.9 | 1.05M | $10→$45Tier caveat | ||
| 02 | Claude Fable 5Anthropic · AvailableLong-running agents · Highest-capability tasks | 9.9editorial | 9.9 | 9.9 | 9.8 | 1M | $10→$50 | ||
| 03 | Claude Opus 5Anthropic · AvailableComplex reasoning · Long-running agents | 9.9editorial | 9.9 | 9.9 | 9.8 | 1M | $5→$25 | ||
| 04 | GPT-5.5OpenAI · AvailableComplex reasoning · Coding | 9.8editorial | 9.8 | 9.8 | 9.7 | 1M | $5→$30 | ||
| 05 | Grok 4.6xAI · AvailableAgentic coding · Knowledge work | 9.8editorial | 9.8 | 9.8 | 9.7 | 500K | $2→$6Tier caveat | ||
| 06 | Claude Opus 4.8Anthropic · AvailableComplex reasoning · Agentic coding | 9.8editorial | 9.8 | 9.8 | 9.6 | 1M | $5→$25 | ||
| 07 | Grok 4.5xAI · AvailableAgentic coding · Knowledge work | 9.7editorial | 9.7 | 9.7 | 9.6 | 500K | $2→$6Tier caveat | ||
| 08 | GPT-5.4OpenAI · AvailableCoding · Agents | 9.7editorial | 9.8 | 9.5 | 9.7 | 1M | $2.5→$15 | ||
| 09 | GPT-5.6 TerraOpenAI · AvailableProduction agents · Coding | 9.6editorial | 9.6 | 9.6 | 9.7 | 1.05M | $4→$18Tier caveat | ||
| 10 | Gemini 3.7 FlashGoogle · AvailableAgentic coding · Multimodal reasoning | 9.6editorial | 9.6 | 9.5 | 9.6 | 1.05M | $0.75→$3.75Tier caveat | ||
| 11 | Claude Sonnet 5Anthropic · AvailableBalanced performance · Production workloads | 9.5editorial | 9.6 | 9.5 | 9.3 | 1M | $2→$10Tier caveat | ||
| 12 | GPT-5.2-CodexOpenAI · AvailableCoding-focused tasks · Type inference | 9.5editorial | 9.7 | 9.3 | 9.4 | 400K | $1.75→$14 | ||
| 13 | Gemini 3.6 FlashGoogle · AvailableAgentic coding · Multimodal tasks | 9.5editorial | 9.5 | 9.4 | 9.5 | 1M | $1.5→$7.5Tier caveat | ||
| 14 | Gemini 3.5 FlashGoogle · AvailableFast multimodal agents · Search grounding | 9.4editorial | 9.4 | 9.4 | 9.4 | 1M | $1.5→$9Tier caveat | ||
| 15 | Kimi K3Moonshot AI · AvailableLong-horizon coding · Knowledge work | 9.3editorial | 9.4 | 9.3 | 9.2 | 1M | $3→$15Tier caveat | ||
| 16 | GLM-5.3Z.ai · AvailableComplex software engineering · Long-context analysis | 9.3editorial | 9.3 | 9.3 | 9.2 | 1M | $1.4→$4.4Tier caveat | ||
| 17 | GPT-5.2OpenAI · AvailableGeneral-purpose · Balanced tasks | 9.2editorial | 9.3 | 9.2 | 9.0 | 400K | $1.75→$14 | ||
| 18 | GPT-OSS-120BOpenAI · AvailableSelf-hosted · Privacy | 9.2editorial | 9.3 | 9.2 | 9.0 | 128K | $0→$0Tier caveat | ||
| 19 | GLM-5Zhipu AI · AvailableBilingual (CN/EN) · Value-focused | 9.2editorial | 9.2 | 9.3 | 9.0 | 200K | $1→$3.2 | ||
| 20 | GPT-5.6 LunaOpenAI · AvailableHigh-volume workflows · Subagents | 9.1editorial | 9.1 | 9.1 | 9.3 | 1.05M | $0.4→$1.8Tier caveat | ||
| 21 | DeepSeek V4 ProDeepSeek · AvailableBudget coding · High-volume | 9.1editorial | 9.1 | 9.2 | 8.9 | 1M | $0.44→$1.32Tier caveat | ||
| 22 | Mistral Medium 3.5Mistral · AvailableEuropean compliance · Agentic coding | 9.0editorial | 9.1 | 9.1 | 8.9 | 256K | $1.5→$7.5 | ||
| 23 | DeepSeek V4 FlashDeepSeek · AvailableLow-cost agent experiments · Responses API | 8.9editorial | 8.9 | 8.9 | 8.9 | 1M | $0.14→$0.28Tier caveat | ||
| 24 | Claude Haiku 4.5Anthropic · AvailableFast responses · High-volume tasks | 8.8editorial | 8.9 | 8.8 | 8.7 | 200K | $1→$5 | ||
| 25 | Gemini 3.5 Flash-LiteGoogle · AvailableHigh-volume automation · Subagents | 8.8editorial | 8.8 | 8.7 | 8.9 | 1M | $0.3→$2.5Tier caveat |
No models match these filters.
Reset the filters or broaden the workload.
Score methodology
Orientation, not manufactured certainty.
The editorial fit score weights coding at 40%, reasoning at 35%, and tool use at 25%. It is a compact decision aid built from the dimensions tracked in this catalog—not a claim that one model is universally best.
Prices show input then output cost per one million tokens. Provider tiers, caching, batches, and hosts can change the actual bill. Open the evidence links and caveat labels before committing to a production choice.
Inspect the source log and review policy →