Every frontier model, benchmarked and put side by side.
A single reference for every major AI product and model on the market — flagship LLMs, image and video generators, and coding agents — with model-to-model comparisons, abstract-reasoning (AGI-track) scores, and a dedicated cybersecurity posture index. Refreshed monthly.
Current model to watch
The featured position is based on the strongest currently available composite evidence, not a claim that one model is universally best for every workload.
Who’s building what
Eight labs currently define the frontier. Each ships more than a chatbot — reasoning models, coding agents, image and video generators, and voice systems all sit under one roof.
Beyond the chatbot
The same labs compete across five other product categories. This isn’t a benchmarked ranking — just a map of who ships what today.
Every tracked model
Filter by lab or search by name. Context window and pricing shown per million tokens where publicly listed.
| Model | Provider | Released | Context | Price in/out ($/M) | ARC‑AGI‑2 | GPQA | SWE‑bench | Best for |
|---|
Latest flagship, per lab
The newest top-tier release from each major provider, right now.
Compare up to three models
Pick any models from the dropdowns below to see full specs, pricing, and benchmark scores side by side.
Abstract reasoning & general intelligence
No model has passed a real AGI test — there isn’t a certified one. These are the closest public proxies: novel visual puzzles, graduate-level science questions, and aggregate “intelligence index” scores that combine dozens of evals.
ARC‑AGI‑2 — novel abstract reasoning
higher = better · % tasks solvedGrid-puzzle tasks designed to resist memorization; average untrained human scores ~66%. Considered the hardest widely-used public reasoning benchmark. Source: arcprize.org public leaderboard.
GPQA Diamond — graduate-level science reasoning
higher = better · % correctPhD-level, Google-proof multiple-choice questions across biology, chemistry and physics.
Artificial-Analysis-style Intelligence Index
composite of ~10 public evals, normalized 0–100A composite capability signal can be useful for orientation, but it is not a definition of general intelligence. Composite rankings should be read alongside coding, agentic, reasoning, cost and deployment evidence.
Cybersecurity posture index
How current models behave as coding assistants and how they resist misuse for offensive cyber tasks.
| Model | Secure-coding tendency | Prompt-injection resistance | Cyberattack-request refusal | Offensive-cyber uplift gating | Notes |
|---|
How often the model’s generated code avoids known insecure patterns (CWE-mapped), per CyberSecEval-style static analysis.
Resistance to hidden instructions embedded in tool outputs, documents, or web content hijacking the model’s behavior.
Whether the lab applies extra restrictions/monitoring to outputs that could materially assist real-world cyberattacks, per its published safety framework.
How this page is built
Public sources only
Every figure traces to a lab’s own release notes/system cards, or a recognized public leaderboard (ARC Prize, GPQA, SWE-bench, Artificial Analysis).
Refreshed monthly
New releases, price changes and re-benchmarked scores are folded in on a monthly pass — labs currently ship a new model every few weeks.
No single “best”
Rankings shift by task. This page favors side-by-side specs over crowning one universal winner.
Directional cyber ratings
Cybersecurity ratings are synthesized qualitative judgments, not one standardized benchmark — treat them as a starting point for your own evaluation.
