AI models ranked
Curated rankings of frontier and open-source AI models for developers. Scores are calibrated against the current frontier so older generations do not look artificially competitive.
58 models - page 1 of 5
OpenAI's GPT-5.6 flagship tier for frontier reasoning, long-horizon coding agents, cybersecurity analysis, biology workflows, and the hardest knowledge work. Sol introduces deeper reasoning controls, including max effort and an ultra mode that can coordinate subagents for complex work.
OpenAI's balanced GPT-5.6 tier for everyday agentic coding, product engineering, analysis, tool workflows, and production assistants. Terra is positioned near GPT-5.5 capability while cutting the token price in half, making it the default choice when Sol's deepest reasoning is not required.
Anthropic's current Opus-tier model for complex reasoning, agentic coding, and high-autonomy workflows. It keeps the strong Claude coding profile with a 1M-token context window and a lower price than Fable 5.
Anthropic's most capable widely released Claude model, built for demanding reasoning and long-horizon agentic work. Use it when autonomy, context depth, and complex multi-step execution matter more than low latency.
Google's stable Gemini 3.5 production model for sustained frontier performance with strong coding, agentic loops, grounding, tool use, and multimodal inputs. It balances capability, 1M-token context, and production cost better than older Gemini 2.5 entries.
xAI's newest flagship chat model with configurable reasoning, agentic tool calling, low hallucination positioning, image input, and a 1M-token context window. Use server-side search tools when current events or live data matter.
OpenAI's newest flagship model for agentic coding, professional knowledge work, data analysis, computer use, and long-running tool workflows. It is positioned as a step up from GPT-5.4 with stronger system understanding, better debugging behavior, and API availability for production developers.
OpenAI's fast, lowest-cost GPT-5.6 tier for high-volume assistance, routing, extraction, quick code edits, support automation, and latency-sensitive workflows. Luna gives teams a current-generation model for scale while preserving stronger reasoning than older budget models.
Google's preview Pro model optimized for software engineering behavior, tool use, agentic workflows, stronger thinking, and grounded multimodal reasoning. Use it for evaluation and advanced workflows where preview volatility is acceptable.
DeepSeek's V4 Pro open-weight model for agentic coding, math, STEM reasoning, and 1M-context workflows. It supports OpenAI-compatible and Anthropic-compatible APIs with thinking and non-thinking modes.
Anthropic's current frontier model. State-of-the-art on SWE-bench, best-in-class instruction following, and extended thinking built-in. The go-to for agentic coding workflows.