Rankings with the tradeoffs left in.
Scan the full catalog, narrow it to your workload, and compare up to three models. Specifications are source-reviewed; fit scores are editorial and their weighting is explicit.
Compare models
Showing 38 models
| Compare | # | Model | Editorial fit | Coding | Reasoning | Tool use | Context | Price / 1M tokens | Evidence |
|---|---|---|---|---|---|---|---|---|---|
| 01 | GPT-6 AstraOpenAI · AvailableHard end-to-end agent work · Computer use | 10.0editorial | 10.0 | 10.0 | 10.0 | 1.05M | $10→$50Tier caveat | ||
| 02 | Claude Fable 5.1Anthropic · AvailableDemanding reasoning · Long-horizon agents | 9.9editorial | 9.9 | 10.0 | 9.9 | 1M | $10→$50Tier caveat | ||
| 03 | GPT-5.6 SolOpenAI · AvailableComplex production workflows · Coding | 9.9editorial | 9.9 | 9.9 | 9.9 | 1.05M | $4→$20Tier caveat | ||
| 04 | Claude Fable 5Anthropic · AvailableLong-running agents · Highest-capability tasks | 9.9editorial | 9.9 | 9.9 | 9.8 | 1M | $10→$50 | ||
| 05 | Claude Opus 5Anthropic · AvailableComplex reasoning · Long-running agents | 9.9editorial | 9.9 | 9.9 | 9.8 | 1M | $5→$25 | ||
| 06 | GPT-5.5OpenAI · AvailableComplex reasoning · Coding | 9.8editorial | 9.8 | 9.8 | 9.7 | 1M | $5→$30 | ||
| 07 | Grok 4.6xAI · AvailableAgentic coding · Knowledge work | 9.8editorial | 9.8 | 9.8 | 9.7 | 500K | $2→$6Tier caveat | ||
| 08 | Claude Opus 4.8Anthropic · AvailableComplex reasoning · Agentic coding | 9.8editorial | 9.8 | 9.8 | 9.6 | 1M | $5→$25 | ||
| 09 | Grok 4.5xAI · AvailableAgentic coding · Knowledge work | 9.7editorial | 9.7 | 9.7 | 9.6 | 500K | $2→$6Tier caveat | ||
| 10 | GPT-5.4OpenAI · AvailableCoding · Agents | 9.7editorial | 9.8 | 9.5 | 9.7 | 1M | $2.5→$15 | ||
| 11 | Gemini 3.8 FlashGoogle · AvailableLong-horizon software engineering · Autonomous agents | 9.7editorial | 9.7 | 9.6 | 9.7 | 1.05M | $0.75→$3.75Tier caveat | ||
| 12 | GPT-5.6 TerraOpenAI · AvailableProduction agents · Coding | 9.6editorial | 9.6 | 9.6 | 9.7 | 1.05M | $2→$12Tier caveat | ||
| 13 | Gemini 3.7 FlashGoogle · AvailableAgentic coding · Multimodal reasoning | 9.6editorial | 9.6 | 9.5 | 9.6 | 1.05M | $0.75→$3.75Tier caveat | ||
| 14 | Qwen3.8-MaxAlibaba / Qwen · AvailableLong-horizon coding · Complex multimodal agents | 9.5editorial | 9.5 | 9.5 | 9.5 | 1M | N/ATier caveat | ||
| 15 | Claude Sonnet 5Anthropic · AvailableBalanced performance · Production workloads | 9.5editorial | 9.6 | 9.5 | 9.3 | 1M | $2→$10Tier caveat | ||
| 16 | GPT-5.2-CodexOpenAI · AvailableCoding-focused tasks · Type inference | 9.5editorial | 9.7 | 9.3 | 9.4 | 400K | $1.75→$14 | ||
| 17 | Gemini 3.6 FlashGoogle · AvailableAgentic coding · Multimodal tasks | 9.5editorial | 9.5 | 9.4 | 9.5 | 1M | $1.5→$7.5Tier caveat | ||
| 18 | Gemini 3.5 FlashGoogle · AvailableFast multimodal agents · Search grounding | 9.4editorial | 9.4 | 9.4 | 9.4 | 1M | $1.5→$9Tier caveat | ||
| 19 | Kimi K3Moonshot AI · AvailableLong-horizon coding · Knowledge work | 9.3editorial | 9.4 | 9.3 | 9.2 | 1M | $3→$15Tier caveat | ||
| 20 | GLM-5.3Z.ai · AvailableComplex software engineering · Long-context analysis | 9.3editorial | 9.3 | 9.3 | 9.2 | 1M | $1.4→$4.4Tier caveat | ||
| 21 | GPT-5.2OpenAI · AvailableGeneral-purpose · Balanced tasks | 9.2editorial | 9.3 | 9.2 | 9.0 | 400K | $1.75→$14 | ||
| 22 | GPT-OSS-120BOpenAI · AvailableSelf-hosted · Privacy | 9.2editorial | 9.3 | 9.2 | 9.0 | 128K | $0→$0Tier caveat | ||
| 23 | GLM-5Zhipu AI · AvailableBilingual (CN/EN) · Value-focused | 9.2editorial | 9.2 | 9.3 | 9.0 | 200K | $1→$3.2 | ||
| 24 | GPT-5.6 LunaOpenAI · AvailableHigh-volume workflows · Subagents | 9.1editorial | 9.1 | 9.1 | 9.3 | 1.05M | $0.2→$1.2Tier caveat | ||
| 25 | GLM-5.3-FlashZ.ai · AvailableCost-sensitive multimodal coding · Long-context agents | 9.1editorial | 9.2 | 9.1 | 9.1 | 1M | $0.15→$0.5Tier caveat | ||
| 26 | DeepSeek V4.1 FlashDeepSeek · AvailableLow-cost multimodal agents · High-throughput coding | 9.1editorial | 9.2 | 9.1 | 9.1 | 1M | $0.15→$0.6Tier caveat | ||
| 27 | Kimi K2.7 CodeMoonshot AI · AvailableLong-horizon software engineering · Coding agents | 9.1editorial | 9.2 | 9.0 | 9.2 | 262K | $0.95→$4Tier caveat | ||
| 28 | DeepSeek V4 ProDeepSeek · AvailableDeepSeek agent workloads · Long-context text workflows | 9.1editorial | 9.1 | 9.2 | 8.9 | 1M | $0.66→$1.98Tier caveat | ||
| 29 | Mistral Medium 3.5Mistral · AvailableEuropean compliance · Agentic coding | 9.0editorial | 9.1 | 9.1 | 8.9 | 256K | $1.5→$7.5 | ||
| 30 | Qwen3.8-FlashAlibaba / Qwen · AvailableHigh-volume multimodal agents · Long-context coding | 9.0editorial | 9.0 | 8.9 | 9.0 | 1M | $0.15→$0.47Tier caveat | ||
| 31 | Qwen3.8-Flash-NextAlibaba / Qwen · AvailableSelf-hosted multimodal agents · Cost-efficient long-context inference | 8.9editorial | 8.9 | 8.8 | 8.9 | 256K | N/ATier caveat | ||
| 32 | Claude Haiku 4.5Anthropic · AvailableFast responses · High-volume tasks | 8.8editorial | 8.9 | 8.8 | 8.7 | 200K | $1→$5 | ||
| 33 | Gemini 3.5 Flash-LiteGoogle · AvailableHigh-volume automation · Subagents | 8.8editorial | 8.8 | 8.7 | 8.9 | 1M | $0.3→$2.5Tier caveat | ||
| 34 | GPT-Live-1OpenAI · AvailableFull-duplex voice agents · Telephony | 1.0editorial | 1.0 | 1.0 | 1.0 | Not stated | N/ATier caveat | ||
| 35 | North Small TranslateCohere · Research/non-commercialMachine translation · Private translation deployments | 1.0editorial | 1.0 | 1.0 | 1.0 | 16K | N/ATier caveat | ||
| — | Gemini 3.8 LiveGoogle · AvailableLow-latency voice agents · Real-time dialogue | N/A | N/A | N/A | N/A | 131K | $0.75→$4.5Tier caveat | ||
| — | Gemini 3.8 Live Extended ThinkingGoogle · AvailableComplex voice agents · Multi-step voice workflows | N/A | N/A | N/A | N/A | 131K | $0.75→$4.5Tier caveat | ||
| — | Qwen3.8-Omni-FlashAlibaba / Qwen · AvailableAudio and video understanding · Multimodal content analysis | N/A | N/A | N/A | N/A | 1M | N/ATier caveat |
No models match these filters.
Reset the filters or broaden the workload.
Orientation, not manufactured certainty.
The editorial fit score weights coding at 40%, reasoning at 35%, and tool use at 25%. It is a compact decision aid built from the dimensions tracked in this catalog—not a claim that one model is universally best. Specialized models that do not map to those dimensions are listed but marked N/A and excluded from editorial ranking.
Prices show input then output cost per one million tokens. Provider tiers, caching, batches, and hosts can change the actual bill. Open the evidence links and caveat labels before committing to a production choice.
Inspect the source log and review policy →