Long-context guide
Best Long-Context Models
For giant documents, repos, transcripts, and agent memory buffers, context window only helps when it is paired with retrieval discipline and pricing you can actually afford.
| Model | Context | Best for | Input $/M |
|---|---|---|---|
| GPT-6 Astra OpenAI | 1.05M | Hard end-to-end agent work | $10 |
| GPT-5.6 Sol OpenAI | 1.05M | Complex production workflows | $4 |
| GPT-5.6 Terra OpenAI | 1.05M | Production agents | $2 |
| GPT-5.6 Luna OpenAI | 1.05M | High-volume workflows | $0.2 |
| Gemini 3.8 Flash | 1.05M | Long-horizon software engineering | $0.75 |
| Gemini 3.7 Flash | 1.05M | Agentic coding | $0.75 |
| GPT-5.5 OpenAI | 1M | Complex reasoning | $5 |
| Claude Fable 5.1 Anthropic | 1M | Demanding reasoning | $10 |
| Claude Fable 5 Anthropic | 1M | Long-running agents | $10 |
| Claude Opus 5 Anthropic | 1M | Complex reasoning | $5 |
| Claude Opus 4.8 Anthropic | 1M | Complex reasoning | $5 |
| GPT-5.4 OpenAI | 1M | Coding | $2.5 |
| Gemini 3.5 Flash | 1M | Fast multimodal agents | $1.5 |
| Gemini 3.6 Flash | 1M | Agentic coding | $1.5 |
| Gemini 3.5 Flash-Lite | 1M | High-volume automation | $0.3 |
| Claude Sonnet 5 Anthropic | 1M | Balanced performance | $2 |
| GLM-5.3-Flash Z.ai | 1M | Cost-sensitive multimodal coding | $0.15 |
| GLM-5.3 Z.ai | 1M | Complex software engineering | $1.4 |
| DeepSeek V4 Pro DeepSeek | 1M | Budget coding | $0.44 |
| DeepSeek V4 Flash DeepSeek | 1M | Low-cost agent experiments | $0.14 |
| Kimi K3 Moonshot AI | 1M | Long-horizon coding | $3 |
| Grok 4.5 xAI | 500K | Agentic coding | $2 |
| Grok 4.6 xAI | 500K | Agentic coding | $2 |
| GPT-5.2-Codex OpenAI | 400K | Coding-focused tasks | $1.75 |
| GPT-5.2 OpenAI | 400K | General-purpose | $1.75 |
| Mistral Medium 3.5 Mistral | 256K | European compliance | $1.5 |
| Claude Haiku 4.5 Anthropic | 200K | Fast responses | $1 |
| GLM-5 Zhipu AI | 200K | Bilingual (CN/EN) | $1 |
| GPT-OSS-120B OpenAI | 128K | Self-hosted | $0 |
| North Small Translate Cohere | 16K | Machine translation | $ |
Fast picks
- GPT-5.5 for hard reasoning over large inputs with a 1M-token context window
- Claude Opus 4.8 for long-running agentic coding and professional document workflows
- Gemini 3.8 Flash for 1,048,576-token, search-grounded long-context multimodal work
- Gemini 3.7 Flash as a stable alternative with the same 1.05M-class context limit
- DeepSeek V4 Pro when 1M context and low token cost matter more than frontier polish