Every frontier model, benchmarked and put side by side.
A single reference for every major AI product and model on the market — flagship LLMs, image and video generators, and coding agents — with model-to-model comparisons, abstract-reasoning (AGI-track) scores, and a dedicated cybersecurity posture index. Refreshed monthly.
Who’s building what
Eight labs currently define the frontier. Each ships more than a chatbot — reasoning models, coding agents, image and video generators, and voice systems all sit under one roof.
Beyond the chatbot
The same labs compete across five other product categories. This isn’t a benchmarked ranking — just a map of who ships what today.
Every tracked model
Filter by lab or search by name. Context window and pricing shown per million tokens where publicly listed. Scores marked EST are our estimate where a lab hasn’t published a directly comparable figure.
| Model | Provider | Released | Context | Price in/out ($/M) | ARC‑AGI‑2 | GPQA | SWE‑bench | Best for |
|---|---|---|---|---|---|---|---|---|
| GPT-6 AstraNEWGATEDFlagship — Critical cyber tier | OpenAI | Sep 3, 2026 | 1.05M | $10 / $50 | 95% | 96% | 74.1% EST | Computer use, agentic software engineering, science |
| Gemini 3.8 FlashNEWLatest stable (Flash) | Google DeepMind | Sep 2, 2026 | 1M | $0.5 / $3 | 86% EST | 94% EST | 73.7% | Fast multimodal & agentic workflows |
| Muse Spark 1.3NEWConsumer assistant | Meta | Sep 2, 2026 | 256K | — / — | 42.5% | 55% EST | 75.4% | Consumer assistant experiences, edges frontier labs on some agentic-coding evals |
| Claude Fable 5.1NEWMythos-class flagship | Anthropic | Sep 1, 2026 | 1M | $15 / $75 | 90% | 92.6% | 81.2% | Frontier reasoning, long-running agents |
| DeepSeek‑V4‑ProFlagship | DeepSeek | Aug 13, 2026 | 128K | $0.5 / $1.5 | 55% EST | 78% EST | 52% EST | Cost-efficient coding & reasoning at scale |
| Grok 4.6NEWFlagship | xAI | Aug 12, 2026 | 500K | $2 / $6 | 58% EST | 89% EST | 65.9% | Chat + coding with real-time X data, long-running agents |
| Qwen3.8‑MaxFlagship | Alibaba | Aug 3, 2026 | 256K | $1.2 / $3.6 | 52% EST | 76% EST | 48% EST | Multilingual + open-weight coding |
| Claude Opus 5Flagship (public) | Anthropic | Jul 24, 2026 | 1M | $15 / $75 | 90.4% | 91.5% | 78% | Professional coding, enterprise work |
| GPT-5.6 SolFlagship reasoning | OpenAI | Jul 9, 2026 | 400K | $35 / $35 | 92.5% | 94.6% | 73% EST | Hardest mixed reasoning/business tasks |
| GPT-5.6 TerraBalanced workhorse | OpenAI | Jul 9, 2026 | 400K | $14 / $14 | 83.9% | 90.5% EST | 70% EST | Default general-purpose OpenAI model |
| GPT-5.6 LunaFast / cost-efficient | OpenAI | Jul 9, 2026 | 128K | $2 / $8 | 59.5% | 82% EST | 55% EST | High-throughput, cost-sensitive apps |
| Grok 4.5Prior flagship | xAI | Jul 8, 2026 | 500K | $2 / $6 | 52.6% | 87% EST | 75% | Multi-agent collaboration tasks |
| Claude Sonnet 5Everyday flagship | Anthropic | Jun 30, 2026 | 1M | $3 / $15 | 60% EST | 89% EST | 72.2% | Day-to-day Claude deployments, chat + agents |
| Gemini 3.5 FlashMid-tier | Google DeepMind | Jun 20, 2026 | 1M | $0.4 / $2.5 | 72.1% | 88% EST | 37% | Balanced speed/cost multimodal tasks |
| Mistral Medium 3.5Flagship | Mistral AI | Apr 28, 2026 | 128K | $1 / $3 | 38% EST | 72% EST | 33% EST | EU data-residency, on-prem deployment |
| GPT-5.5Prior flagship | OpenAI | Apr 23, 2026 | 400K | $35 / $35 | 85% | 90% EST | 58.6% | General knowledge work |
| Gemini 3.1 ProFrontier value | Google DeepMind | Feb 19, 2026 | 1M–2M | $2 / $12 | 77.1% | 94.3% | 63.8% | Best reasoning-per-dollar, science QA, video |
| Claude Haiku 4.5Fast / low-cost | Anthropic | Oct 15, 2025 | 200K | $1 / $5 | 25% EST | 73% | 73.3% | High-volume, latency-sensitive tasks |
| Llama 4 MaverickOpen-weight flagship | Meta | Apr 5, 2025 | 1M | $0.2 / $0.6 | 18% EST | 48% EST | 28% EST | Self-hosted multimodal deployment |
| Llama 4 ScoutOpen-weight, long-context | Meta | Apr 5, 2025 | 10M | $0.15 / $0.4 | 14% EST | 44% EST | 22% EST | Extreme-long-context open workloads |
Latest flagship, per lab
The newest top-tier release from each major provider, right now.
GPT-6 Astra
Gemini 3.8 Flash
Muse Spark 1.3
Claude Fable 5.1
Grok 4.6
Compare up to three models
Pick any models from the dropdowns below to see full specs, pricing, and benchmark scores side by side.
GPT-6 Astra
GPT-5.6 Sol
Gemini 3.8 Flash
Abstract reasoning & general intelligence
No model has passed a real AGI test — there isn’t a certified one. These are the closest public proxies: novel visual puzzles, graduate-level science questions, and aggregate “intelligence index” scores that combine dozens of evals.
ARC‑AGI‑2 — novel abstract reasoning
higher = better · % tasks solvedGrid-puzzle tasks designed to resist memorization; average untrained human scores ~66%. Considered the hardest widely-used public reasoning benchmark. Source: arcprize.org public leaderboard.
GPQA Diamond — graduate-level science reasoning
higher = better · % correctPhD-level, Google-proof multiple-choice questions across biology, chemistry and physics.
Artificial-Analysis-style Intelligence Index
composite of ~10 public evals, normalized 0–100A weighted blend of reasoning, knowledge, coding and math benchmarks used as a rough single-number stand-in for general capability.
Cybersecurity posture index
How current models behave as coding assistants and how they resist misuse for offensive cyber tasks.
| Model | Secure-coding tendency | Prompt-injection resistance | Cyberattack-request refusal | Uplift gating | Notes |
|---|---|---|---|---|---|
| GPT-6 Astra | High | Medium | High | GATED | First model OpenAI has designated “Critical” (its highest tier) under the Preparedness Framework for cyber risk. Public rollout ships refusing advanced offensive-cyber tasks; full capability limited to vetted defenders in OpenAI’s Daybreak program. Scored 100% on OpenAI’s internal ExploitBench. |
| Claude Opus 5 / Fable 5.1 | High | High | High | GATED | Anthropic’s Responsible Scaling Policy applies extra cyber-uplift safeguards at this capability tier; Mythos variant relaxes some restrictions for vetted orgs only. |
| GPT-5.6 Sol / Terra | High | Medium | High | GATED | OpenAI’s Preparedness Framework gates high-uplift cyber capability; strong on secure-code suggestion benchmarks. |
| Gemini 3.1 Pro / 3.8 Flash | Medium | Medium | High | GATED | Google’s Frontier Safety Framework applies dangerous-capability evaluations pre-release; a dedicated “Flash Cyber” variant ships alongside the base 3.8 Flash model. |
| Grok 4.6 | Medium | Medium | Medium | ungated | Lighter published safety-framework detail than the other three frontier labs; independent red-team coverage is thinner. |
| Llama 4 (open-weight) | Medium | Lower | Lower | ungated | Open weights mean any safety fine-tuning can be stripped by a downstream deployer — protections aren’t guaranteed at inference time. |
| DeepSeek‑V4‑Pro | Medium | Lower | Lower | ungated | Limited independent third-party security red-teaming publicly available as of this update. |
How often the model’s generated code avoids known insecure patterns (CWE-mapped), per CyberSecEval-style static analysis.
Resistance to hidden instructions embedded in tool outputs, documents, or web content hijacking the model’s behavior.
Whether the lab applies extra restrictions/monitoring to outputs that could materially assist real-world cyberattacks, per its published safety framework.
How this page is built
Public sources only
Every figure traces to a lab’s own release notes/system cards, or a recognized public leaderboard (ARC Prize, GPQA, SWE-bench, Artificial Analysis).
Refreshed monthly
New releases, price changes and re-benchmarked scores are folded in on a monthly pass — labs currently ship a new model every few weeks.
No single “best”
Rankings shift by task. This page favors side-by-side specs over crowning one universal winner.
Estimated fields, marked
Smaller or less-benchmarked providers don’t always publish ARC-AGI-2/GPQA/SWE-bench figures directly. Where we’ve interpolated from a closely related published eval, the directory marks the model with an EST badge.
