The Intelligence Index — Frontier AI Models, Benchmarks & Safety Intelligence
    EVERY MAJOR AI LAB · ONE LIVE REFERENCE

    Every frontier model, benchmarked and put side by side.

    A single reference for every major AI product and model on the market — flagship LLMs, image and video generators, and coding agents — with model-to-model comparisons, abstract-reasoning (AGI-track) scores, and a dedicated cybersecurity posture index. Refreshed monthly.

    models tracked across labs
    current featured frontier model
    top public ARC-AGI-2 score (abstract reasoning)
    02 — LABS & PRODUCTS

    Who’s building what

    Eight labs currently define the frontier. Each ships more than a chatbot — reasoning models, coding agents, image and video generators, and voice systems all sit under one roof.

    03 — PRODUCT CATEGORIES

    Beyond the chatbot

    The same labs compete across five other product categories. This isn’t a benchmarked ranking — just a map of who ships what today.

    04 — DIRECTORY

    Every tracked model

    Filter by lab or search by name. Context window and pricing shown per million tokens where publicly listed.

    ModelProviderReleasedContext Price in/out ($/M)ARC‑AGI‑2GPQASWE‑benchBest for
    No models match that filter.
    05 — THIS MONTH

    Latest flagship, per lab

    The newest top-tier release from each major provider, right now.

    06 — HEAD TO HEAD

    Compare up to three models

    Pick any models from the dropdowns below to see full specs, pricing, and benchmark scores side by side.

    07 — AGI TRACK

    Abstract reasoning & general intelligence

    No model has passed a real AGI test — there isn’t a certified one. These are the closest public proxies: novel visual puzzles, graduate-level science questions, and aggregate “intelligence index” scores that combine dozens of evals.

    Evidence rule: Provider-reported frontier results are retained as provider evidence and are not automatically blended into independently comparable rankings. Benchmark versions, evaluation environments and tool configurations can materially change results.

    ARC‑AGI‑2 — novel abstract reasoning

    higher = better · % tasks solved

    Grid-puzzle tasks designed to resist memorization; average untrained human scores ~66%. Considered the hardest widely-used public reasoning benchmark. Source: arcprize.org public leaderboard.

    GPQA Diamond — graduate-level science reasoning

    higher = better · % correct

    PhD-level, Google-proof multiple-choice questions across biology, chemistry and physics.

    Artificial-Analysis-style Intelligence Index

    composite of ~10 public evals, normalized 0–100

    A composite capability signal can be useful for orientation, but it is not a definition of general intelligence. Composite rankings should be read alongside coding, agentic, reasoning, cost and deployment evidence.

    08 — CYBERSECURITY TRACK

    Cybersecurity posture index

    How current models behave as coding assistants and how they resist misuse for offensive cyber tasks.

    Reading this table: there is no single certified “AI cybersecurity leaderboard” — labs publish results using different internal red-team suites and methodologies (Meta’s CyberSecEval, SecBench, and lab-specific system-card evaluations). Ratings below are a directional synthesis of published safety/system-card disclosures and third-party security research, not a precise head-to-head score. Industry-wide baseline from CyberSecEval: LLMs suggest insecure code in roughly 1 of every 3 completions on average, and comply with cyberattack-assistance requests roughly half the time absent added safeguards — frontier labs now layer additional filtering on top of the base model to push these numbers down.
    ModelSecure-coding tendencyPrompt-injection resistanceCyberattack-request refusalOffensive-cyber uplift gatingNotes
    Secure-coding tendency

    How often the model’s generated code avoids known insecure patterns (CWE-mapped), per CyberSecEval-style static analysis.

    Prompt-injection resistance

    Resistance to hidden instructions embedded in tool outputs, documents, or web content hijacking the model’s behavior.

    Offensive-cyber uplift gating

    Whether the lab applies extra restrictions/monitoring to outputs that could materially assist real-world cyberattacks, per its published safety framework.

    09 — METHODOLOGY

    How this page is built

    01
    Public sources only

    Every figure traces to a lab’s own release notes/system cards, or a recognized public leaderboard (ARC Prize, GPQA, SWE-bench, Artificial Analysis).

    02
    Refreshed monthly

    New releases, price changes and re-benchmarked scores are folded in on a monthly pass — labs currently ship a new model every few weeks.

    03
    No single “best”

    Rankings shift by task. This page favors side-by-side specs over crowning one universal winner.

    04
    Directional cyber ratings

    Cybersecurity ratings are synthesized qualitative judgments, not one standardized benchmark — treat them as a starting point for your own evaluation.

    Independent reference page. Benchmark figures are classified by evidence type and should be verified against primary sources before procurement or production deployment. Not affiliated with OpenAI, Anthropic, Google, xAI, Meta, Mistral, DeepSeek or Alibaba. Trademarks belong to their respective owners. Figures are best-effort as of the “updated” date above and may lag official leaderboards.