OpenAI current flagship โ Sol/Terra/Luna tiers, 1.05M context
GA since July 2026 and now the OpenAI default. Sol is the flagship tier (88.8%, 91.9% on Sol Ultra), Terra a balanced mid-tier at roughly half the cost, Luna the cheap high-volume option โ same 1.05M-token context and 128K output across all three. Ahead of GPT-5.5 and Claude Opus 4.8 on OpenAI's own benchmark, still behind the gated Claude Mythos 5. Terra is the sane default for most agentic and coding work; reach for Sol only when you are actually bottlenecked on quality.
OpenAI's previous flagship โ superseded by GPT-5.6
Still a capable model and still served, but GPT-5.6 Terra beats it on quality at a similar price. No longer a reason to build new on plain GPT-5 โ migrate to GPT-5.6 (Terra for most workloads, Sol if you need the ceiling).
OpenAI's cost-efficient GPT-5-tier model โ most of the quality at a fraction of the cost
The default now for classification, extraction, summarisation, and simple chat โ meaningfully sharper than GPT-4o mini for the same rough price. Inherits GPT-5's unified reasoning at lower latency. First choice for cost-sensitive apps that outgrew GPT-3.5-era models.
OpenAI previous-gen multimodal model โ text, vision, audio
Two generations back from GPT-5.6, and still widely deployed โ the API is stable and a huge number of production RAG and agentic pipelines are built on it. JSON mode, function calling, and vision are battle-tested. Fine to keep running if it is already tuned and evaluated; start new projects on GPT-5.6 instead.
OpenAI's previous-gen cost-efficient model โ superseded by GPT-5 mini
Still cheap and still works, but GPT-5 mini beats it on quality at a comparable price โ there is no longer a strong reason to pick this for a new build. Reasonable to leave alone in an existing pipeline that already meets its bar.
OpenAI's first open-weight model โ 120B params, fully downloadable
Historic moment โ OpenAI releasing open weights. Competitive with Llama 4 and Mistral Large. Fine-tunable and self-hostable. Great for teams wanting OpenAI quality with full control.
Anthropic's most intelligent generally available model โ Mythos-class reasoning
The new Anthropic flagship. Shares its underlying model with the gated Mythos 5 tier, minus the dual-use safety loosening โ for everyday production work the distinction rarely matters. Extended thinking, mature MCP tool use, and the deepest long-context reasoning Anthropic has shipped. Default choice for architecture review and hard multi-step agent work.
Frontier-class agentic coding and computer use, at unchanged Opus pricing
Shipped July 24, 2026 at the same $5/$25 as Opus 4.8, but more than doubles it on Frontier-Bench and triples the next-best model on ARC-AGI 3. Lands within half a percent of Fable 5's peak coding score at roughly a third of the cost. If your workload is agentic coding or computer use specifically, this is now the better cost-to-quality pick over Fable 5 โ reach for Fable 5 instead when the task is open-ended reasoning rather than execution.
Anthropic's production workhorse โ Fable-class reasoning at Sonnet latency and cost
The one most teams should actually be running in production. Meaningfully sharper than Claude 4 Sonnet on coding and long-context retrieval, without the latency and cost of the Fable tier. MCP tool ecosystem is mature. First pick for agentic workflows, code generation, and day-to-day engineering assistance.
Anthropic's fast tier โ Sonnet 4-level quality, Haiku pricing
Missing from this catalog until now, which undersold it โ this is Anthropic's first Haiku with extended thinking and computer use, matching Claude 4 Sonnet on reasoning and coding at $1/$5 per million tokens. 73.3% on SWE-bench Verified is a real number for a "fast" model. Drop-in replacement for Haiku 3.5 or Sonnet 4 wherever you were paying more than you needed to.
Anthropic's previous-gen workhorse โ still solid, no longer the default
One generation back from Sonnet 5, and still a perfectly good model โ the gap is real but not dramatic. Reasonable choice if you have existing evals tuned against it. New projects should start on Sonnet 5 instead.
Anthropic's older workhorse with 200K context and computer use
Two generations behind Sonnet 5 now. Computer use API still works and 200K context is still respectable, but there is no real reason to start a new project here โ treat it as a fallback for legacy integrations you have not migrated yet.
Google's most intelligent workhorse model yet โ for coding and agents
Launched August 13, 2026, and the sharpest thing Google ships right now โ including against its own Pro tier, which has been frozen at 3.1 since February. Lets you dial thinking level up or down per request instead of picking a fixed model for latency vs. quality. If you are choosing a Gemini model today for coding or agentic work, start here, not with "Pro."
Google's Pro-tier model โ capable, but its own Flash line has passed it
Still the current "Pro" release, but Google has not updated this tier since February 2026 while Flash raced to 3.7 โ and 3.5 Flash already outperforms 3.1 Pro on agentic and coding benchmarks per Google's own comparisons. Native 2M context and deep multimodal understanding still make it the right call for very long documents. For coding and agents specifically, pick Gemini 3.7 Flash instead.
Google's previous Pro-tier model โ superseded by Gemini 3.1 Pro
A full major version behind the current Pro tier now. The 2M context window and multimodal handling were the headline then; Gemini 3.1 Pro does the same job with a newer model underneath. Migrate new work to 3.1 Pro, or to 3.7 Flash if the workload is coding or agentic.
Google's previous speed-optimised model โ superseded by Gemini 3.7 Flash
Two Flash generations behind now. 3.7 Flash costs about the same, uses meaningfully fewer tokens per task, and adds adjustable thinking levels this version never had. No reason to start new work on 2.0 Flash.
Meta's new proprietary flagship โ closed-weight, agentic-first
Meta's Superintelligence Labs shipped this as a genuinely new direction, not a Llama update โ closed-weight, built explicitly for multi-agent orchestration (as either the planning agent or a delegated subagent), with a 1M-token context window and text/image/video/audio/PDF input. The 1.2 update (Aug 5, 2026) added a dedicated coding focus plus Muse Code, a terminal agent. The real tradeoff: you lose everything that made Llama useful for self-hosting โ no local deployment, no fine-tuning, no community variants. Pick this for Meta's agentic capability specifically; pick Llama if open weights still matter to you.
Meta's flagship open model โ 400B MoE with 128 experts
Massive leap for open-source. 400B MoE architecture with 128 experts runs efficiently on 8ร H100. Matches frontier closed models on most benchmarks. Best open model for enterprise self-hosting โ and now the more relevant Meta model for that use case, since Meta's newer Muse Spark line is closed-weight.
Meta's proven open model โ battle-tested at 70B params
Battle-tested in thousands of production deployments. Runs on a single A100 80GB. Excellent for self-hosted RAG, fine-tuning, and cost-sensitive pipelines. Huge ecosystem of fine-tunes.
1.6T MoE, 1M context, MIT-licensed โ DeepSeek's new flagship
Released April 24, 2026, unifying what used to be separate R1 (reasoning) and V3 (general) models into one, with three selectable reasoning modes (Non-Think, Think High, Think Max). 1.6T total / 49B active parameters, a genuinely usable 1M-token context at roughly 27% of the FLOPs and 10% of the KV-cache footprint of the previous generation, MIT-licensed. At $0.435/$0.87 per million tokens it is dramatically cheaper than any closed frontier model at a comparable level โ still the best self-hostable option for reasoning-heavy work.
The efficient V4 variant โ 284B MoE, same 1M context, MIT license
Same architecture and 1M context as V4 Pro at a fifth of the size (284B total / 13B active) and roughly a fifth of the price ($0.14/$0.28 per million tokens). Give up some ceiling on the hardest reasoning tasks in exchange for materially cheaper, faster inference โ the right default for high-volume pipelines that do not need V4 Pro's full capability.
Open-weight chain-of-thought reasoning model โ superseded by DeepSeek V4
Made frontier reasoning accessible to everyone when it shipped, and still works โ but DeepSeek V4 folded R1's reasoning specialism and V3's general capability into a single model with selectable thinking modes, at a fraction of the inference cost. Migrate to DeepSeek V4 Pro; there is no longer a reason to run R1 and V3 as separate models.
685B MoE model trained for $5.5M โ superseded by DeepSeek V4
Proved you can train frontier models affordably, and the efficiency story continues in V4, which replaces it directly with a larger context window, lower inference cost per the DeepSeek team's own figures, and unified reasoning modes. No reason to start new work on V3.
Alibaba's 2.4T multimodal flagship โ text, image, and video in one model
Launched August 3, 2026 (2.4T total / 95B active MoE, ~1M-token context, native text/image/video input) at $2/$6 per million tokens. A real step up from the Qwen 3.5 line, not just a version bump โ this is now Alibaba's serious multimodal-reasoning play, not just its best multilingual model. Best choice in this catalog for CJK-heavy or multimodal workloads that also need frontier-adjacent reasoning.
Alibaba's previous multilingual model โ superseded by Qwen 3.8 Max
Still fine for lightweight multilingual work at 9Bโ72B sizes with permissive licensing, but Qwen 3.8 Max is a generational jump in both scale and multimodality. Pick 3.8 Max for new projects unless you specifically need Qwen 3.5's smaller, self-hostable sizes.
Google's open model family โ 26B and 31B instruction-tuned
Best small-to-medium open model from Google. 26B A4B variant uses mixture of experts for efficiency. Strong at instruction following and reasoning. Good JAX and Keras ecosystem.
675B open-weight MoE, Apache 2.0 โ the largest open model from a major lab
675B total / 41B active MoE, 262K context, fully Apache 2.0 โ download it, fine-tune it, run it on your own hardware, no license negotiation. At $0.50/$1.50 per million tokens via Mistral's own API it is also cheap if you would rather not self-host. The strongest fully-open option in this catalog right now for teams that need both frontier-adjacent quality and real data-residency control.
Previous Mistral flagship โ superseded by Mistral Large 3
Was the best EU-hosted option with data-residency guarantees; Mistral Large 3 replaces it directly and adds a full open-weight Apache 2.0 release on top, so the "EU option" argument no longer requires giving up self-hosting. Migrate new work to Large 3.
Microsoft's 14B model punching above its weight on reasoning
Remarkable reasoning for its size โ beats many 70B models on math and logic benchmarks. Runs on consumer GPUs. Ideal for edge deployment and latency-sensitive applications.
xAI's current model โ built for long-running agents and coding
Released August 12, 2026, an agent- and coding-focused upgrade over Grok 4.5 rather than a context or pricing change โ same 500K context, same $2/$6 short-context rate. The real change is sustained multi-step task performance. Watch the long-context pricing: past 200K tokens, input, cached input, and output all double for the entire request, not just the overage.
xAI's previous model โ superseded by Grok 4.6
Two releases behind now (4.5, then 4.6). Still functional via the API, but there is no remaining reason to build new on Grok 4 specifically โ go straight to 4.6.
Zhipu AI's 754B frontier model from China's leading AI lab
One of the largest dense models available. Strong Chinese language capabilities and general reasoning. Open weights on Hugging Face. Interesting for multilingual and research applications.
Cohere's enterprise model optimised for RAG and tool use
Purpose-built for enterprise RAG. Excellent citation generation and grounding. Multilingual at 10 languages. Cohere Coral SDK simplifies integration. Good for accuracy-critical search applications.
OpenAI's original cost-efficient chat model โ now superseded
Fully superseded by GPT-4o mini at similar cost and much higher quality. Avoid for new projects. Migrate to gpt-4o-mini or a modern open-weight alternative.