Model Comparison
This page puts the currently active models side by side so you can pick one at a glance. Scan the table for context window, input modalities, what each model is best for, and its native API protocol, then read the “How to choose” notes below for quick guidance. Prices are not shown here because they change; live pricing and availability are on the model marketplace in the dashboard, listed per model.
Comparison table
Anthropic (native protocol: Anthropic)
| Model | Context window | Modalities | Best for | Native protocol |
|---|---|---|---|---|
| Claude Fable 5 | 1,000,000 tokens | Text, image | Frontier autonomous coding and deep knowledge work; always-on adaptive thinking | Anthropic |
| Claude Sonnet 5 | 1,000,000 tokens | Text, image | Balanced default for production coding, high-concurrency services, long-context agents | Anthropic |
| Claude Opus 4.8 | 1,000,000 tokens | Text, image | Previous-gen frontier for autonomous agents and long-horizon work; still served | Anthropic |
| Claude Haiku 4.5 | 200,000 tokens | Text, image | Fast, low-cost, high-volume real-time tasks and summarization | Anthropic |
OpenAI (native protocol: OpenAI)
| Model | Context window | Modalities | Best for | Native protocol |
|---|---|---|---|---|
| GPT-5.6 Sol | 1,050,000 tokens | Text, image, PDF | Flagship deep reasoning for complex coding, analysis, and planning | OpenAI |
| GPT-5.6 Terra | 1,050,000 tokens | Text, image, PDF | Balanced, mid-range price for general coding and business apps | OpenAI |
| GPT-5.6 Luna | 1,050,000 tokens | Text, image | Cost-effective, cost-first for high-frequency light tasks and real-time | OpenAI |
| GPT-5.5 | 1,050,000 tokens | Text, image | Previous-gen flagship for multi-step autonomous tasks and research | OpenAI |
Google (native protocol: OpenAI)
| Model | Context window | Modalities | Best for | Native protocol |
|---|---|---|---|---|
| Gemini 3.1 Pro | 1,050,000 tokens | Text, image, PDF | Flagship Pro tier for advanced reasoning, coding, and multimodal analysis | OpenAI |
| Gemini 3.5 Flash | 1,100,000 tokens | Text, image, video | Workhorse for high-throughput coding and agentic execution at Flash-tier speed | OpenAI |
| Gemini 3 Flash | 1,050,000 tokens | Text, image | Lightweight, low-latency, cost-sensitive tasks | OpenAI |
Open / China (native protocol: OpenAI)
| Model | Context window | Modalities | Best for | Native protocol |
|---|---|---|---|---|
| GLM-5.2 | 1,000,000 tokens | Text, image, video | Top-tier open-weights model for long-horizon engineering and Chinese scenarios | OpenAI |
| GLM-5.3 | 1,000,000 tokens | Text | New-gen flagship, reasoning always on, for complex software engineering and long-horizon agents | OpenAI |
| GLM-5.1 | 128,000 tokens | Text | Previous-gen open-source flagship for agentic engineering; still served | OpenAI |
How to choose
- Balanced default: Start with Claude Sonnet 5 for most production coding and agentic work, or GPT-5.6 Terra, the balanced mid-range option on the OpenAI side.
- Strongest reasoning: For the most complex coding and deep knowledge work, choose Claude Fable 5 or GPT-5.6 Sol.
- Cheapest and fastest: For high-frequency, latency-sensitive tasks, use Claude Haiku 4.5, GPT-5.6 Luna, or Gemini 3 Flash.
- Longest context: Gemini 3.5 Flash offers the largest window at 1,100,000 tokens; several others reach 1,000,000 to 1,050,000. If you need long context on open weights, use GLM-5.2.
- Video input: Choose Gemini 3.5 Flash or GLM-5.2, which accept video natively.
Full reference
See the full Models Reference for every model, including more models still coming online (such as video and additional open-weights models). Live pricing and current availability are listed there and on the model marketplace in the dashboard.
Last updated on