AI Wisdom
๐Ÿง 

Text Generation & Reasoning

Foundation models, fine-tunes, and inference APIs powering modern AI applications.

Production ยท 17Stable ยท 7Experimental ยท 1Superseded ยท 934 total
โ† All categories

GPT-5.6

Production
5/5

OpenAI current flagship โ€” Sol/Terra/Luna tiers, 1.05M context

GA since July 2026 and now the OpenAI default. Sol is the flagship tier (88.8%, 91.9% on Sol Ultra), Terra a balanced mid-tier at roughly half the cost, Luna the cheap high-volume option โ€” same 1.05M-token context and 128K output across all three. Ahead of GPT-5.5 and Claude Opus 4.8 on OpenAI's own benchmark, still behind the gated Claude Mythos 5. Terra is the sane default for most agentic and coding work; reach for Sol only when you are actually bottlenecked on quality.

Proprietary

GPT-5

Superseded
4/5

OpenAI's previous flagship โ€” superseded by GPT-5.6

Still a capable model and still served, but GPT-5.6 Terra beats it on quality at a similar price. No longer a reason to build new on plain GPT-5 โ€” migrate to GPT-5.6 (Terra for most workloads, Sol if you need the ceiling).

Proprietary

GPT-5 mini

Production
4/5

OpenAI's cost-efficient GPT-5-tier model โ€” most of the quality at a fraction of the cost

The default now for classification, extraction, summarisation, and simple chat โ€” meaningfully sharper than GPT-4o mini for the same rough price. Inherits GPT-5's unified reasoning at lower latency. First choice for cost-sensitive apps that outgrew GPT-3.5-era models.

Proprietary

GPT-4o

Production
4/5

OpenAI previous-gen multimodal model โ€” text, vision, audio

Two generations back from GPT-5.6, and still widely deployed โ€” the API is stable and a huge number of production RAG and agentic pipelines are built on it. JSON mode, function calling, and vision are battle-tested. Fine to keep running if it is already tuned and evaluated; start new projects on GPT-5.6 instead.

3/5

OpenAI's previous-gen cost-efficient model โ€” superseded by GPT-5 mini

Still cheap and still works, but GPT-5 mini beats it on quality at a comparable price โ€” there is no longer a strong reason to pick this for a new build. Reasonable to leave alone in an existing pipeline that already meets its bar.

Proprietary
4/5

OpenAI's first open-weight model โ€” 120B params, fully downloadable

Historic moment โ€” OpenAI releasing open weights. Competitive with Llama 4 and Mistral Large. Fine-tunable and self-hostable. Great for teams wanting OpenAI quality with full control.

Open Source

Claude Fable 5

Production
5/5

Anthropic's most intelligent generally available model โ€” Mythos-class reasoning

The new Anthropic flagship. Shares its underlying model with the gated Mythos 5 tier, minus the dual-use safety loosening โ€” for everyday production work the distinction rarely matters. Extended thinking, mature MCP tool use, and the deepest long-context reasoning Anthropic has shipped. Default choice for architecture review and hard multi-step agent work.

Proprietary

Claude Opus 5

Production
5/5

Frontier-class agentic coding and computer use, at unchanged Opus pricing

Shipped July 24, 2026 at the same $5/$25 as Opus 4.8, but more than doubles it on Frontier-Bench and triples the next-best model on ARC-AGI 3. Lands within half a percent of Fable 5's peak coding score at roughly a third of the cost. If your workload is agentic coding or computer use specifically, this is now the better cost-to-quality pick over Fable 5 โ€” reach for Fable 5 instead when the task is open-ended reasoning rather than execution.

Proprietary

Claude Sonnet 5

Production
5/5

Anthropic's production workhorse โ€” Fable-class reasoning at Sonnet latency and cost

The one most teams should actually be running in production. Meaningfully sharper than Claude 4 Sonnet on coding and long-context retrieval, without the latency and cost of the Fable tier. MCP tool ecosystem is mature. First pick for agentic workflows, code generation, and day-to-day engineering assistance.

4/5

Anthropic's fast tier โ€” Sonnet 4-level quality, Haiku pricing

Missing from this catalog until now, which undersold it โ€” this is Anthropic's first Haiku with extended thinking and computer use, matching Claude 4 Sonnet on reasoning and coding at $1/$5 per million tokens. 73.3% on SWE-bench Verified is a real number for a "fast" model. Drop-in replacement for Haiku 3.5 or Sonnet 4 wherever you were paying more than you needed to.

Proprietary

Claude 4 Sonnet

Production
4/5

Anthropic's previous-gen workhorse โ€” still solid, no longer the default

One generation back from Sonnet 5, and still a perfectly good model โ€” the gap is real but not dramatic. Reasonable choice if you have existing evals tuned against it. New projects should start on Sonnet 5 instead.

Proprietary
3/5

Anthropic's older workhorse with 200K context and computer use

Two generations behind Sonnet 5 now. Computer use API still works and 200K context is still respectable, but there is no real reason to start a new project here โ€” treat it as a fallback for legacy integrations you have not migrated yet.

Proprietary
5/5

Google's most intelligent workhorse model yet โ€” for coding and agents

Launched August 13, 2026, and the sharpest thing Google ships right now โ€” including against its own Pro tier, which has been frozen at 3.1 since February. Lets you dial thinking level up or down per request instead of picking a fixed model for latency vs. quality. If you are choosing a Gemini model today for coding or agentic work, start here, not with "Pro."

Proprietary

Gemini 3.1 Pro

Production
4/5

Google's Pro-tier model โ€” capable, but its own Flash line has passed it

Still the current "Pro" release, but Google has not updated this tier since February 2026 while Flash raced to 3.7 โ€” and 3.5 Flash already outperforms 3.1 Pro on agentic and coding benchmarks per Google's own comparisons. Native 2M context and deep multimodal understanding still make it the right call for very long documents. For coding and agents specifically, pick Gemini 3.7 Flash instead.

Proprietary

Gemini 2.5 Pro

Superseded
3/5

Google's previous Pro-tier model โ€” superseded by Gemini 3.1 Pro

A full major version behind the current Pro tier now. The 2M context window and multimodal handling were the headline then; Gemini 3.1 Pro does the same job with a newer model underneath. Migrate new work to 3.1 Pro, or to 3.7 Flash if the workload is coding or agentic.

Proprietary
3/5

Google's previous speed-optimised model โ€” superseded by Gemini 3.7 Flash

Two Flash generations behind now. 3.7 Flash costs about the same, uses meaningfully fewer tokens per task, and adds adjustable thinking levels this version never had. No reason to start new work on 2.0 Flash.

Proprietary

Muse Spark

Production
5/5

Meta's new proprietary flagship โ€” closed-weight, agentic-first

Meta's Superintelligence Labs shipped this as a genuinely new direction, not a Llama update โ€” closed-weight, built explicitly for multi-agent orchestration (as either the planning agent or a delegated subagent), with a 1M-token context window and text/image/video/audio/PDF input. The 1.2 update (Aug 5, 2026) added a dedicated coding focus plus Muse Code, a terminal agent. The real tradeoff: you lose everything that made Llama useful for self-hosting โ€” no local deployment, no fine-tuning, no community variants. Pick this for Meta's agentic capability specifically; pick Llama if open weights still matter to you.

Proprietary
5/5

Meta's flagship open model โ€” 400B MoE with 128 experts

Massive leap for open-source. 400B MoE architecture with 128 experts runs efficiently on 8ร— H100. Matches frontier closed models on most benchmarks. Best open model for enterprise self-hosting โ€” and now the more relevant Meta model for that use case, since Meta's newer Muse Spark line is closed-weight.

Open Source

Llama 3.3 70B

Production
4/5

Meta's proven open model โ€” battle-tested at 70B params

Battle-tested in thousands of production deployments. Runs on a single A100 80GB. Excellent for self-hosted RAG, fine-tuning, and cost-sensitive pipelines. Huge ecosystem of fine-tunes.

Open Source

DeepSeek V4 Pro

Production
5/5

1.6T MoE, 1M context, MIT-licensed โ€” DeepSeek's new flagship

Released April 24, 2026, unifying what used to be separate R1 (reasoning) and V3 (general) models into one, with three selectable reasoning modes (Non-Think, Think High, Think Max). 1.6T total / 49B active parameters, a genuinely usable 1M-token context at roughly 27% of the FLOPs and 10% of the KV-cache footprint of the previous generation, MIT-licensed. At $0.435/$0.87 per million tokens it is dramatically cheaper than any closed frontier model at a comparable level โ€” still the best self-hostable option for reasoning-heavy work.

Open Source
4/5

The efficient V4 variant โ€” 284B MoE, same 1M context, MIT license

Same architecture and 1M context as V4 Pro at a fifth of the size (284B total / 13B active) and roughly a fifth of the price ($0.14/$0.28 per million tokens). Give up some ceiling on the hardest reasoning tasks in exchange for materially cheaper, faster inference โ€” the right default for high-volume pipelines that do not need V4 Pro's full capability.

Open Source

DeepSeek R1

Superseded
3/5

Open-weight chain-of-thought reasoning model โ€” superseded by DeepSeek V4

Made frontier reasoning accessible to everyone when it shipped, and still works โ€” but DeepSeek V4 folded R1's reasoning specialism and V3's general capability into a single model with selectable thinking modes, at a fraction of the inference cost. Migrate to DeepSeek V4 Pro; there is no longer a reason to run R1 and V3 as separate models.

Open Source

DeepSeek V3

Superseded
3/5

685B MoE model trained for $5.5M โ€” superseded by DeepSeek V4

Proved you can train frontier models affordably, and the efficiency story continues in V4, which replaces it directly with a larger context window, lower inference cost per the DeepSeek team's own figures, and unified reasoning modes. No reason to start new work on V3.

Open Source

Qwen 3.8 Max

Production
5/5

Alibaba's 2.4T multimodal flagship โ€” text, image, and video in one model

Launched August 3, 2026 (2.4T total / 95B active MoE, ~1M-token context, native text/image/video input) at $2/$6 per million tokens. A real step up from the Qwen 3.5 line, not just a version bump โ€” this is now Alibaba's serious multimodal-reasoning play, not just its best multilingual model. Best choice in this catalog for CJK-heavy or multimodal workloads that also need frontier-adjacent reasoning.

Proprietary

Qwen 3.5

Superseded
3/5

Alibaba's previous multilingual model โ€” superseded by Qwen 3.8 Max

Still fine for lightweight multilingual work at 9Bโ€“72B sizes with permissive licensing, but Qwen 3.8 Max is a generational jump in both scale and multimodality. Pick 3.8 Max for new projects unless you specifically need Qwen 3.5's smaller, self-hostable sizes.

Open Source

Gemma 4

Stable
4/5

Google's open model family โ€” 26B and 31B instruction-tuned

Best small-to-medium open model from Google. 26B A4B variant uses mixture of experts for efficiency. Strong at instruction following and reasoning. Good JAX and Keras ecosystem.

Open Source

Mistral Large 3

Production
5/5

675B open-weight MoE, Apache 2.0 โ€” the largest open model from a major lab

675B total / 41B active MoE, 262K context, fully Apache 2.0 โ€” download it, fine-tune it, run it on your own hardware, no license negotiation. At $0.50/$1.50 per million tokens via Mistral's own API it is also cheap if you would rather not self-host. The strongest fully-open option in this catalog right now for teams that need both frontier-adjacent quality and real data-residency control.

Open Source

Mistral Large 2

Superseded
3/5

Previous Mistral flagship โ€” superseded by Mistral Large 3

Was the best EU-hosted option with data-residency guarantees; Mistral Large 3 replaces it directly and adds a full open-weight Apache 2.0 release on top, so the "EU option" argument no longer requires giving up self-hosting. Migrate new work to Large 3.

Proprietary

Phi-4

Stable
4/5

Microsoft's 14B model punching above its weight on reasoning

Remarkable reasoning for its size โ€” beats many 70B models on math and logic benchmarks. Runs on consumer GPUs. Ideal for edge deployment and latency-sensitive applications.

Open Source

Grok 4.6

Production
4/5

xAI's current model โ€” built for long-running agents and coding

Released August 12, 2026, an agent- and coding-focused upgrade over Grok 4.5 rather than a context or pricing change โ€” same 500K context, same $2/$6 short-context rate. The real change is sustained multi-step task performance. Watch the long-context pricing: past 200K tokens, input, cached input, and output all double for the entire request, not just the overage.

Proprietary

Grok 4

Superseded
3/5

xAI's previous model โ€” superseded by Grok 4.6

Two releases behind now (4.5, then 4.6). Still functional via the API, but there is no remaining reason to build new on Grok 4 specifically โ€” go straight to 4.6.

Proprietary

GLM-5

Experimental
4/5

Zhipu AI's 754B frontier model from China's leading AI lab

One of the largest dense models available. Strong Chinese language capabilities and general reasoning. Open weights on Hugging Face. Interesting for multilingual and research applications.

Open Source
3/5

Cohere's enterprise model optimised for RAG and tool use

Purpose-built for enterprise RAG. Excellent citation generation and grounding. Multilingual at 10 languages. Cohere Coral SDK simplifies integration. Good for accuracy-critical search applications.

GPT-3.5 Turbo

Superseded
2/5

OpenAI's original cost-efficient chat model โ€” now superseded

Fully superseded by GPT-4o mini at similar cost and much higher quality. Avoid for new projects. Migrate to gpt-4o-mini or a modern open-weight alternative.

Proprietary