DeepSeek V3
685B MoE model trained for $5.5M — superseded by DeepSeek V4
Verdict
Proved you can train frontier models affordably, and the efficiency story continues in V4, which replaces it directly with a larger context window, lower inference cost per the DeepSeek team's own figures, and unified reasoning modes. No reason to start new work on V3.
Other Text Generation & Reasoning
- GPT-5.6Production
OpenAI current flagship — Sol/Terra/Luna tiers, 1.05M context
- GPT-5Superseded
OpenAI's previous flagship — superseded by GPT-5.6
- GPT-5 miniProduction
OpenAI's cost-efficient GPT-5-tier model — most of the quality at a fraction of the cost
- GPT-4oProduction
OpenAI previous-gen multimodal model — text, vision, audio
- GPT-4o miniStable
OpenAI's previous-gen cost-efficient model — superseded by GPT-5 mini
- GPT-OSS 120BStable
OpenAI's first open-weight model — 120B params, fully downloadable

