The Intelligence Index — AI Models, Benchmarks & Cybersecurity Ratings
    EVERY MAJOR AI LAB · ONE LIVE REFERENCE

    Every frontier model, benchmarked and put side by side.

    A single reference for every major AI product and model on the market — flagship LLMs, image and video generators, and coding agents — with model-to-model comparisons, abstract-reasoning (AGI-track) scores, and a dedicated cybersecurity posture index. Refreshed monthly.

    20
    models tracked across 8 labs
    Claude Fable 5.1
    current #1 on the Intelligence Index
    95%
    top public ARC-AGI-2 score (abstract reasoning)
    OpenAI  GPT-6 Astra  released Sep 3, 2026
    Google DeepMind  Gemini 3.8 Flash  released Sep 2, 2026
    Meta  Muse Spark 1.3  released Sep 2, 2026
    Anthropic  Claude Fable 5.1  released Sep 1, 2026
    DeepSeek  DeepSeek‑V4‑Pro  released Aug 13, 2026
    xAI  Grok 4.6  released Aug 12, 2026
    Alibaba  Qwen3.8‑Max  released Aug 3, 2026
    Anthropic  Claude Opus 5  released Jul 24, 2026
    OpenAI  GPT-5.6 Sol  released Jul 9, 2026
    OpenAI  GPT-5.6 Terra  released Jul 9, 2026
    OpenAI  GPT-6 Astra  released Sep 3, 2026
    Google DeepMind  Gemini 3.8 Flash  released Sep 2, 2026
    Meta  Muse Spark 1.3  released Sep 2, 2026
    Anthropic  Claude Fable 5.1  released Sep 1, 2026
    DeepSeek  DeepSeek‑V4‑Pro  released Aug 13, 2026
    xAI  Grok 4.6  released Aug 12, 2026
    Alibaba  Qwen3.8‑Max  released Aug 3, 2026
    Anthropic  Claude Opus 5  released Jul 24, 2026
    OpenAI  GPT-5.6 Sol  released Jul 9, 2026
    OpenAI  GPT-5.6 Terra  released Jul 9, 2026
    01 — LABS & PRODUCTS

    Who’s building what

    Eight labs currently define the frontier. Each ships more than a chatbot — reasoning models, coding agents, image and video generators, and voice systems all sit under one roof.

    AN
    Anthropic
    Claude Fable 5.1
    Claude family — Haiku, Sonnet, Opus, plus the new Mythos-tier (Fable/Mythos 5.1) above Opus.
    ChatReasoningCoding agents
    OP
    OpenAI
    GPT-6 Astra
    GPT-6 Astra (Sept 3, 2026) tops the GPT-5.6 (Sol/Terra/Luna) family, with computer use and a Critical-tier cybersecurity designation, alongside Sora video and native image generation.
    ChatReasoningImageVideo
    GO
    Google DeepMind
    Gemini 3.8 Flash
    Gemini 3.x line spanning Flash to Pro/Deep Think, plus Veo video and Imagen.
    ChatReasoningVideoMultimodal
    XA
    xAI
    Grok 4.6
    Grok 4.x line, now under SpaceXAI, with deep X/real-time data integration and long-running agentic workflows.
    ChatReal-time dataAgents
    ME
    Meta
    Llama 4 Maverick
    Open-weight Llama 4 line plus the newer Muse Spark consumer assistant models.
    Open-weightChatSelf-host
    DE
    DeepSeek
    DeepSeek‑V4‑Pro
    Cost-efficient reasoning and coding models, popular for API-heavy workloads.
    CodingValueOpen-weight
    MI
    Mistral AI
    Mistral Medium 3.5
    EU-based models built for data residency, on-prem deployment and multilingual use.
    EU/on-premMultilingual
    AL
    Alibaba
    Qwen3.8‑Max
    Qwen line — strong multilingual and coding performance, widely used open-weight base.
    Open-weightMultilingualCoding
    02 — PRODUCT CATEGORIES

    Beyond the chatbot

    The same labs compete across five other product categories. This isn’t a benchmarked ranking — just a map of who ships what today.

    Image generationtext/image → image
    GPT Image (OpenAI)Imagen (Google)Midjourney v7Flux (Black Forest Labs)
    Video generationtext/image → video
    Sora 2 (OpenAI)Veo 3 (Google)Runway Gen-4Kling / Seedream (ByteDance)
    Voice & audiospeech, music, cloning
    ElevenLabsGemini Live AudioOpenAI Realtime Voice
    Coding agentsIDE / CLI / autonomous
    Claude CodeGitHub Copilot / CodexGemini CLICursor
    Search & research agentsmulti-step web research
    Claude Cowork / ResearchChatGPT Deep ResearchGemini Deep ResearchGrok DeepSearch
    03 — DIRECTORY

    Every tracked model

    Filter by lab or search by name. Context window and pricing shown per million tokens where publicly listed. Scores marked EST are our estimate where a lab hasn’t published a directly comparable figure.

    ModelProviderReleasedContext Price in/out ($/M)ARC‑AGI‑2GPQASWE‑benchBest for
    GPT-6 AstraNEWGATEDFlagship — Critical cyber tier OpenAI Sep 3, 2026 1.05M $10 / $50 95% 96% 74.1% EST Computer use, agentic software engineering, science
    Gemini 3.8 FlashNEWLatest stable (Flash) Google DeepMind Sep 2, 2026 1M $0.5 / $3 86% EST 94% EST 73.7% Fast multimodal & agentic workflows
    Muse Spark 1.3NEWConsumer assistant Meta Sep 2, 2026 256K — / — 42.5% 55% EST 75.4% Consumer assistant experiences, edges frontier labs on some agentic-coding evals
    Claude Fable 5.1NEWMythos-class flagship Anthropic Sep 1, 2026 1M $15 / $75 90% 92.6% 81.2% Frontier reasoning, long-running agents
    DeepSeek‑V4‑ProFlagship DeepSeek Aug 13, 2026 128K $0.5 / $1.5 55% EST 78% EST 52% EST Cost-efficient coding & reasoning at scale
    Grok 4.6NEWFlagship xAI Aug 12, 2026 500K $2 / $6 58% EST 89% EST 65.9% Chat + coding with real-time X data, long-running agents
    Qwen3.8‑MaxFlagship Alibaba Aug 3, 2026 256K $1.2 / $3.6 52% EST 76% EST 48% EST Multilingual + open-weight coding
    Claude Opus 5Flagship (public) Anthropic Jul 24, 2026 1M $15 / $75 90.4% 91.5% 78% Professional coding, enterprise work
    GPT-5.6 SolFlagship reasoning OpenAI Jul 9, 2026 400K $35 / $35 92.5% 94.6% 73% EST Hardest mixed reasoning/business tasks
    GPT-5.6 TerraBalanced workhorse OpenAI Jul 9, 2026 400K $14 / $14 83.9% 90.5% EST 70% EST Default general-purpose OpenAI model
    GPT-5.6 LunaFast / cost-efficient OpenAI Jul 9, 2026 128K $2 / $8 59.5% 82% EST 55% EST High-throughput, cost-sensitive apps
    Grok 4.5Prior flagship xAI Jul 8, 2026 500K $2 / $6 52.6% 87% EST 75% Multi-agent collaboration tasks
    Claude Sonnet 5Everyday flagship Anthropic Jun 30, 2026 1M $3 / $15 60% EST 89% EST 72.2% Day-to-day Claude deployments, chat + agents
    Gemini 3.5 FlashMid-tier Google DeepMind Jun 20, 2026 1M $0.4 / $2.5 72.1% 88% EST 37% Balanced speed/cost multimodal tasks
    Mistral Medium 3.5Flagship Mistral AI Apr 28, 2026 128K $1 / $3 38% EST 72% EST 33% EST EU data-residency, on-prem deployment
    GPT-5.5Prior flagship OpenAI Apr 23, 2026 400K $35 / $35 85% 90% EST 58.6% General knowledge work
    Gemini 3.1 ProFrontier value Google DeepMind Feb 19, 2026 1M–2M $2 / $12 77.1% 94.3% 63.8% Best reasoning-per-dollar, science QA, video
    Claude Haiku 4.5Fast / low-cost Anthropic Oct 15, 2025 200K $1 / $5 25% EST 73% 73.3% High-volume, latency-sensitive tasks
    Llama 4 MaverickOpen-weight flagship Meta Apr 5, 2025 1M $0.2 / $0.6 18% EST 48% EST 28% EST Self-hosted multimodal deployment
    Llama 4 ScoutOpen-weight, long-context Meta Apr 5, 2025 10M $0.15 / $0.4 14% EST 44% EST 22% EST Extreme-long-context open workloads
    GPT-6 AstraNEWFlagship — Critical cyber tier
    OpenAI
    RELEASED
    Sep 3, 2026
    CONTEXT
    1.05M
    PRICE IN/OUT
    $10 / $50
    ARC-AGI-2
    95%
    GPQA
    96%
    SWE-BENCH
    74.1% EST
    BEST FOR
    Computer use, agentic software engineering, science
    Gemini 3.8 FlashNEWLatest stable (Flash)
    Google DeepMind
    RELEASED
    Sep 2, 2026
    CONTEXT
    1M
    PRICE IN/OUT
    $0.5 / $3
    ARC-AGI-2
    86% EST
    GPQA
    94% EST
    SWE-BENCH
    73.7%
    BEST FOR
    Fast multimodal & agentic workflows
    Muse Spark 1.3NEWConsumer assistant
    Meta
    RELEASED
    Sep 2, 2026
    CONTEXT
    256K
    PRICE IN/OUT
    — / —
    ARC-AGI-2
    42.5%
    GPQA
    55% EST
    SWE-BENCH
    75.4%
    BEST FOR
    Consumer assistant experiences, edges frontier labs on some agentic-coding evals
    Claude Fable 5.1NEWMythos-class flagship
    Anthropic
    RELEASED
    Sep 1, 2026
    CONTEXT
    1M
    PRICE IN/OUT
    $15 / $75
    ARC-AGI-2
    90%
    GPQA
    92.6%
    SWE-BENCH
    81.2%
    BEST FOR
    Frontier reasoning, long-running agents
    DeepSeek‑V4‑ProFlagship
    DeepSeek
    RELEASED
    Aug 13, 2026
    CONTEXT
    128K
    PRICE IN/OUT
    $0.5 / $1.5
    ARC-AGI-2
    55% EST
    GPQA
    78% EST
    SWE-BENCH
    52% EST
    BEST FOR
    Cost-efficient coding & reasoning at scale
    Grok 4.6NEWFlagship
    xAI
    RELEASED
    Aug 12, 2026
    CONTEXT
    500K
    PRICE IN/OUT
    $2 / $6
    ARC-AGI-2
    58% EST
    GPQA
    89% EST
    SWE-BENCH
    65.9%
    BEST FOR
    Chat + coding with real-time X data, long-running agents
    Qwen3.8‑MaxFlagship
    Alibaba
    RELEASED
    Aug 3, 2026
    CONTEXT
    256K
    PRICE IN/OUT
    $1.2 / $3.6
    ARC-AGI-2
    52% EST
    GPQA
    76% EST
    SWE-BENCH
    48% EST
    BEST FOR
    Multilingual + open-weight coding
    Claude Opus 5Flagship (public)
    Anthropic
    RELEASED
    Jul 24, 2026
    CONTEXT
    1M
    PRICE IN/OUT
    $15 / $75
    ARC-AGI-2
    90.4%
    GPQA
    91.5%
    SWE-BENCH
    78%
    BEST FOR
    Professional coding, enterprise work
    GPT-5.6 SolFlagship reasoning
    OpenAI
    RELEASED
    Jul 9, 2026
    CONTEXT
    400K
    PRICE IN/OUT
    $35 / $35
    ARC-AGI-2
    92.5%
    GPQA
    94.6%
    SWE-BENCH
    73% EST
    BEST FOR
    Hardest mixed reasoning/business tasks
    GPT-5.6 TerraBalanced workhorse
    OpenAI
    RELEASED
    Jul 9, 2026
    CONTEXT
    400K
    PRICE IN/OUT
    $14 / $14
    ARC-AGI-2
    83.9%
    GPQA
    90.5% EST
    SWE-BENCH
    70% EST
    BEST FOR
    Default general-purpose OpenAI model
    GPT-5.6 LunaFast / cost-efficient
    OpenAI
    RELEASED
    Jul 9, 2026
    CONTEXT
    128K
    PRICE IN/OUT
    $2 / $8
    ARC-AGI-2
    59.5%
    GPQA
    82% EST
    SWE-BENCH
    55% EST
    BEST FOR
    High-throughput, cost-sensitive apps
    Grok 4.5Prior flagship
    xAI
    RELEASED
    Jul 8, 2026
    CONTEXT
    500K
    PRICE IN/OUT
    $2 / $6
    ARC-AGI-2
    52.6%
    GPQA
    87% EST
    SWE-BENCH
    75%
    BEST FOR
    Multi-agent collaboration tasks
    Claude Sonnet 5Everyday flagship
    Anthropic
    RELEASED
    Jun 30, 2026
    CONTEXT
    1M
    PRICE IN/OUT
    $3 / $15
    ARC-AGI-2
    60% EST
    GPQA
    89% EST
    SWE-BENCH
    72.2%
    BEST FOR
    Day-to-day Claude deployments, chat + agents
    Gemini 3.5 FlashMid-tier
    Google DeepMind
    RELEASED
    Jun 20, 2026
    CONTEXT
    1M
    PRICE IN/OUT
    $0.4 / $2.5
    ARC-AGI-2
    72.1%
    GPQA
    88% EST
    SWE-BENCH
    37%
    BEST FOR
    Balanced speed/cost multimodal tasks
    Mistral Medium 3.5Flagship
    Mistral AI
    RELEASED
    Apr 28, 2026
    CONTEXT
    128K
    PRICE IN/OUT
    $1 / $3
    ARC-AGI-2
    38% EST
    GPQA
    72% EST
    SWE-BENCH
    33% EST
    BEST FOR
    EU data-residency, on-prem deployment
    GPT-5.5Prior flagship
    OpenAI
    RELEASED
    Apr 23, 2026
    CONTEXT
    400K
    PRICE IN/OUT
    $35 / $35
    ARC-AGI-2
    85%
    GPQA
    90% EST
    SWE-BENCH
    58.6%
    BEST FOR
    General knowledge work
    Gemini 3.1 ProFrontier value
    Google DeepMind
    RELEASED
    Feb 19, 2026
    CONTEXT
    1M–2M
    PRICE IN/OUT
    $2 / $12
    ARC-AGI-2
    77.1%
    GPQA
    94.3%
    SWE-BENCH
    63.8%
    BEST FOR
    Best reasoning-per-dollar, science QA, video
    Claude Haiku 4.5Fast / low-cost
    Anthropic
    RELEASED
    Oct 15, 2025
    CONTEXT
    200K
    PRICE IN/OUT
    $1 / $5
    ARC-AGI-2
    25% EST
    GPQA
    73%
    SWE-BENCH
    73.3%
    BEST FOR
    High-volume, latency-sensitive tasks
    Llama 4 MaverickOpen-weight flagship
    Meta
    RELEASED
    Apr 5, 2025
    CONTEXT
    1M
    PRICE IN/OUT
    $0.2 / $0.6
    ARC-AGI-2
    18% EST
    GPQA
    48% EST
    SWE-BENCH
    28% EST
    BEST FOR
    Self-hosted multimodal deployment
    Llama 4 ScoutOpen-weight, long-context
    Meta
    RELEASED
    Apr 5, 2025
    CONTEXT
    10M
    PRICE IN/OUT
    $0.15 / $0.4
    ARC-AGI-2
    14% EST
    GPQA
    44% EST
    SWE-BENCH
    22% EST
    BEST FOR
    Extreme-long-context open workloads
    04 — THIS MONTH

    Latest flagship, per lab

    The newest top-tier release from each major provider, right now.

    GPT-6 Astra

    OpenAI
    Sep 3, 2026
    Computer use, agentic software engineering, science. Flagship — Critical cyber tier.
    95%
    ARC-AGI-2
    1.05M
    CONTEXT
    $10
    $/M IN

    Gemini 3.8 Flash

    Google DeepMind
    Sep 2, 2026
    Fast multimodal & agentic workflows. Latest stable (Flash).
    86% EST
    ARC-AGI-2
    1M
    CONTEXT
    $0.5
    $/M IN

    Muse Spark 1.3

    Meta
    Sep 2, 2026
    Consumer assistant experiences, edges frontier labs on some agentic-coding evals. Consumer assistant.
    42.5%
    ARC-AGI-2
    256K
    CONTEXT
    $/M IN

    Claude Fable 5.1

    Anthropic
    Sep 1, 2026
    Frontier reasoning, long-running agents. Mythos-class flagship.
    90%
    ARC-AGI-2
    1M
    CONTEXT
    $15
    $/M IN

    Grok 4.6

    xAI
    Aug 12, 2026
    Chat + coding with real-time X data, long-running agents. Flagship.
    58% EST
    ARC-AGI-2
    500K
    CONTEXT
    $2
    $/M IN
    05 — HEAD TO HEAD

    Compare up to three models

    Pick any models from the dropdowns below to see full specs, pricing, and benchmark scores side by side.

    OpenAI

    GPT-6 Astra

    TIER
    Flagship — Critical cyber tier
    RELEASED
    Sep 3, 2026
    CONTEXT WINDOW
    1.05M
    PRICE ($/M IN / OUT)
    $10 / $50
    ARC-AGI-2
    95%
    GPQA DIAMOND
    96%
    SWE-BENCH
    74.1% EST
    INTELLIGENCE INDEX
    61
    BEST FOR
    Computer use, agentic software engineering, science
    OpenAI

    GPT-5.6 Sol

    TIER
    Flagship reasoning
    RELEASED
    Jul 9, 2026
    CONTEXT WINDOW
    400K
    PRICE ($/M IN / OUT)
    $35 / $35
    ARC-AGI-2
    92.5%
    GPQA DIAMOND
    94.6%
    SWE-BENCH
    73% EST
    INTELLIGENCE INDEX
    61
    BEST FOR
    Hardest mixed reasoning/business tasks
    Google DeepMind

    Gemini 3.8 Flash

    TIER
    Latest stable (Flash)
    RELEASED
    Sep 2, 2026
    CONTEXT WINDOW
    1M
    PRICE ($/M IN / OUT)
    $0.5 / $3
    ARC-AGI-2
    86% EST
    GPQA DIAMOND
    94% EST
    SWE-BENCH
    73.7%
    INTELLIGENCE INDEX
    58 EST
    BEST FOR
    Fast multimodal & agentic workflows
    06 — AGI TRACK

    Abstract reasoning & general intelligence

    No model has passed a real AGI test — there isn’t a certified one. These are the closest public proxies: novel visual puzzles, graduate-level science questions, and aggregate “intelligence index” scores that combine dozens of evals.

    New this update: GPT-6 Astra (Sept 3, 2026) reports up to 99.9% on ARC-AGI-3 — a harder, newer generation of the puzzle set below, measured on OpenAI’s own stateful adapter harness — and 97.6% on FrontierMath Tier 4. Independent runs on the standard neutral harness put the same ARC-AGI-3 figure closer to 62.7%, and ARC Prize’s own standard-harness ARC-AGI-2 score for Astra is 95.0%, which is on the same scale as the bars below and is now included there. On the independent Artificial Analysis Intelligence Index, Astra scores ~61 — roughly level with GPT-5.6 Sol and behind Claude Fable 5.1 and Meta’s Muse Spark 1.3.

    ARC‑AGI‑2 — novel abstract reasoning

    higher = better · % tasks solved

    Grid-puzzle tasks designed to resist memorization; average untrained human scores ~66%. Considered the hardest widely-used public reasoning benchmark. Source: arcprize.org public leaderboard.

    GPT-6 Astra
    95%
    GPT-5.6 Sol
    92.5%
    Claude Opus 5
    90.4%
    Claude Fable 5.1
    90%
    Gemini 3.8 Flash
    86%
    GPT-5.5
    85%
    GPT-5.6 Terra
    83.9%
    Gemini 3.1 Pro
    77.1%
    Gemini 3.5 Flash
    72.1%
    Human baseline*
    66%
    Grok 4.6
    58%
    Muse Spark 1.3
    42.5%

    GPQA Diamond — graduate-level science reasoning

    higher = better · % correct

    PhD-level, Google-proof multiple-choice questions across biology, chemistry and physics.

    GPT-6 Astra
    96%
    GPT-5.6 Sol
    94.6%
    Gemini 3.1 Pro
    94.3%
    Gemini 3.8 Flash
    94%
    Claude Fable 5.1
    92.6%
    Claude Opus 5
    91.5%
    GPT-5.6 Terra
    90.5%
    Claude Sonnet 5
    89%
    Grok 4.6
    89%

    Artificial-Analysis-style Intelligence Index

    composite of ~10 public evals, normalized 0–100

    A weighted blend of reasoning, knowledge, coding and math benchmarks used as a rough single-number stand-in for general capability.

    Claude Fable 5.1
    66%
    Claude Opus 5
    63%
    Muse Spark 1.3
    62%
    GPT-6 Astra
    61%
    Grok 4.6
    61%
    GPT-5.6 Sol
    61%
    Gemini 3.8 Flash
    58%
    Gemini 3.1 Pro
    57%
    Claude Sonnet 5
    53%
    07 — CYBERSECURITY TRACK

    Cybersecurity posture index

    How current models behave as coding assistants and how they resist misuse for offensive cyber tasks.

    Reading this table: there is no single certified “AI cybersecurity leaderboard” — labs publish results using different internal red-team suites and methodologies (Meta’s CyberSecEval, SecBench, and lab-specific system-card evaluations). Ratings below are a directional synthesis of published safety/system-card disclosures and third-party security research, not a precise head-to-head score. Industry-wide baseline from CyberSecEval: LLMs suggest insecure code in roughly 1 of every 3 completions on average, and comply with cyberattack-assistance requests roughly half the time absent added safeguards — frontier labs now layer additional filtering on top of the base model to push these numbers down.
    ModelSecure-coding tendencyPrompt-injection resistanceCyberattack-request refusalUplift gatingNotes
    GPT-6 Astra High Medium High GATED First model OpenAI has designated “Critical” (its highest tier) under the Preparedness Framework for cyber risk. Public rollout ships refusing advanced offensive-cyber tasks; full capability limited to vetted defenders in OpenAI’s Daybreak program. Scored 100% on OpenAI’s internal ExploitBench.
    Claude Opus 5 / Fable 5.1 High High High GATED Anthropic’s Responsible Scaling Policy applies extra cyber-uplift safeguards at this capability tier; Mythos variant relaxes some restrictions for vetted orgs only.
    GPT-5.6 Sol / Terra High Medium High GATED OpenAI’s Preparedness Framework gates high-uplift cyber capability; strong on secure-code suggestion benchmarks.
    Gemini 3.1 Pro / 3.8 Flash Medium Medium High GATED Google’s Frontier Safety Framework applies dangerous-capability evaluations pre-release; a dedicated “Flash Cyber” variant ships alongside the base 3.8 Flash model.
    Grok 4.6 Medium Medium Medium ungated Lighter published safety-framework detail than the other three frontier labs; independent red-team coverage is thinner.
    Llama 4 (open-weight) Medium Lower Lower ungated Open weights mean any safety fine-tuning can be stripped by a downstream deployer — protections aren’t guaranteed at inference time.
    DeepSeek‑V4‑Pro Medium Lower Lower ungated Limited independent third-party security red-teaming publicly available as of this update.
    Secure-coding tendency

    How often the model’s generated code avoids known insecure patterns (CWE-mapped), per CyberSecEval-style static analysis.

    Prompt-injection resistance

    Resistance to hidden instructions embedded in tool outputs, documents, or web content hijacking the model’s behavior.

    Offensive-cyber uplift gating

    Whether the lab applies extra restrictions/monitoring to outputs that could materially assist real-world cyberattacks, per its published safety framework.

    08 — METHODOLOGY

    How this page is built

    01
    Public sources only

    Every figure traces to a lab’s own release notes/system cards, or a recognized public leaderboard (ARC Prize, GPQA, SWE-bench, Artificial Analysis).

    02
    Refreshed monthly

    New releases, price changes and re-benchmarked scores are folded in on a monthly pass — labs currently ship a new model every few weeks.

    03
    No single “best”

    Rankings shift by task. This page favors side-by-side specs over crowning one universal winner.

    04
    Estimated fields, marked

    Smaller or less-benchmarked providers don’t always publish ARC-AGI-2/GPQA/SWE-bench figures directly. Where we’ve interpolated from a closely related published eval, the directory marks the model with an EST badge.

    · arcprize.org — ARC-AGI-2 public leaderboard· Artificial Analysis — Intelligence Index· OpenAI — GPT-6 Astra system card & launch benchmarks· xAI — Grok 4.6 announcement· Meta AI — CyberSecEval / Purple Llama benchmark suite· Provider system cards & release notes (OpenAI, Anthropic, Google DeepMind, xAI, Meta)· SWE-bench Verified / SWE-bench Pro public leaderboards· GPQA Diamond public leaderboard· Independent trackers (DataLearner, BenchLM, BenchmarkList) for cross-model comparison

    Independent reference page. Not affiliated with OpenAI, Anthropic, Google, xAI, Meta, Mistral, DeepSeek or Alibaba. Trademarks belong to their respective owners. Figures are best-effort as of the “updated” date above and may lag official leaderboards.

    Omkar Nath Nandi

    Omkar Nath Nandi

    17+ Years in Full Stack Marketing. AI Assisted Marketing Strategist. Built 200+ AI Assisted Marketing Tools. Specialist in Product Marketing, SaaS, B2B, B2C, SEO and Performance Marketing. Trained 100,000+ Professionals. IIT and IIM Guest Faculty | IIM Calcutta Alumni | Ex-Entrepreneur.

    17 Years in Digital Marketing | 12 Years as Trainer

    Creator • Builder • Analyzer • Marketer • Writer

    Currently leading global digital marketing initiatives for a Gartner SIEM Magic Quadrant Leader. Over the past 17+ years, I have partnered with 500+ businesses to drive measurable growth through SEO, Performance Marketing, Product Marketing, SaaS Marketing, Demand Generation, and AI Assisted Marketing. I have built 200+ AI Assisted Marketing Tools and trained more than 100,000 professionals while helping organizations accelerate pipeline growth, strengthen brand visibility, and deliver measurable business outcomes.

    1,00,000+
    Students
    4M+
    Quora Views
    1M+
    Blog Views
    Guest Faculty
    IIT & IIM
    $2M+ Managed Ad Budget
    200+ Websites Built
    200+ Vibe-Coded Apps

    AI & Full Stack Tools Repositories

    Cybersecurity Threat AI Toolkit ↗

    Advanced AI-assisted toolset engineered for cyber threat analysis and real-time security mapping.

    Digital Marketing & Utility Engine ↗