AI Model Timeline

Tracking the accelerating release frequency of frontier AI models: every milestone model from OpenAI, Google, Anthropic, Meta, Mistral, xAI, DeepSeek, Alibaba and others, with benchmark scores, API pricing and who led the field when.

Live Status: Data through 2026-09-10
Cloud · Frontier Models

Macro Overview

World Models · no metric
2019
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
2020
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
2021
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
2022
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
2023
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
OpenAI
GPT-2
OpenAI
2019-02-14
1.5B
GPT-3
OpenAI
2020-06-11
175B
DALL-E
OpenAI
2021-01-05
DALL-E 2
OpenAI
2022-04-06
ChatGPT
OpenAI
2022-11-30
GPT-4
OpenAI
2023-03-14
DALL-E 3
OpenAI
2023-10-03
GPT-4 Turbo
OpenAI
2023-11-06
Sora
OpenAI
2024-02-15
GPT-4o
OpenAI
2024-05-13
OpenAI o1
OpenAI
2024-09-12
Operator
OpenAI
2025-01-23
o3-mini
OpenAI
2025-01-31
GPT-4.5
OpenAI
2025-02-27
o3
OpenAI
2025-04-16
o4-mini
OpenAI
2025-04-16
GPT-5-codex
OpenAI
2025-09-15
ChatGPT 5.1
OpenAI
2025-11-12
GPT-5.1-Codex-Max
OpenAI
2025-11-19
GPT-5.3 Codex
OpenAI
2026-02-05
Sora 2
OpenAI
2025-09-15
GPT-5.3 Instant
OpenAI
2026-03-05
GPT-5.4 Thinking
OpenAI
2026-03-05
GPT-5.4 Pro
OpenAI
2026-03-05
GPT-5.4 mini
OpenAI
2026-03-17
GPT-5.4 nano
OpenAI
2026-03-17
GPT-5.5
OpenAI
2026-04-23
ChatGPT Images 2.0
OpenAI
2026-04-21
GPT Image 1.5
OpenAI
2025-12-16
GPT-5.5 Instant
OpenAI
2026-05-05
GPT-5.6 Sol
OpenAI
2026-06-26
GPT-5.6 Terra
OpenAI
2026-06-26
GPT-5.6 Luna
OpenAI
2026-06-26
GPT-6 Astra
OpenAI
2026-09-03
Google
Bard
Google
2023-03-21
Gemini 1.0
Google
2023-12-06
Gemini 1.5 Pro
Google
2024-02-15
Gemini 1.5 Flash
Google
2024-05-14
Imagen 3
Google
2024-08-28
Gemini 2.0 Flash
Google
2025-02-05
Gemini 2.5
Google
2025-06-17
Gemini 2.5 Pro
Google
2025-03-25
Gemini 3
Google
2025-11-18
Gemini 3 Pro Image Preview
Google
2025-11-20
Gemini 3.1 Pro
Google
2026-02-19
Veo
Google
2024-05-14
Veo 2
Google
2024-12-16
Veo 3
Google
2025-05-13
Veo 3.1
Google
2025-10-15
Nano Banana 2
Google
2026-02-26
Gemini 3.1 Flash Image
Google
2026-02-26
Gemini 3.5 Flash
Google
2026-05-19
Gemini 3 Flash
Google
2025-12-17
Nano Banana 2 Lite
Google
2026-06-30
Gemini Omni Flash
Google
2026-06-30
Gemini 3.6 Flash
Google
2026-07-21
Gemini 3.7 Flash
Google
2026-08-13
Gemini 3.5 Flash-Lite
Google
2026-07-21
Gemini 3.5 Flash Cyber
Google
2026-07-21
Gemini 3.8 Flash
Google
2026-09-02
Gemini 3.8 Flash Cyber
Google
2026-09-02
Gemini Omni 1.1 Flash
Google
2026-08-27
Gemini Robotics 1.5
Google
2025-09-25
Gemini Robotics 2
Google
2026-07-30
Genie 3
Google
2025-08-05
Anthropic
Claude 1
Anthropic
2023-03-14
Claude 2
Anthropic
2023-07-11
Claude 3 Opus
Anthropic
2024-03-04
Claude 3.5 Sonnet
Anthropic
2024-06-21
Claude 3.5 Sonnet v2
Anthropic
2024-10-22
Claude 3.5 Haiku
Anthropic
2024-10-22
Claude 3.7 Sonnet
Anthropic
2025-02-24
Claude Sonnet 4
Anthropic
2025-05-22
Claude Opus 4
Anthropic
2025-05-22
Claude Opus 4.1
Anthropic
2025-08-05
Claude Sonnet 4.5
Anthropic
2025-09-01
Claude Haiku 4.5
Anthropic
2025-10-01
Claude Opus 4.5
Anthropic
2025-11-24
Claude Opus 4.6
Anthropic
2026-02-05
Claude Sonnet 4.6
Anthropic
2026-02-17
Claude Mythos Preview
Anthropic
2026-04-07
Claude Opus 4.7
Anthropic
2026-04-16
Claude Opus 4.8
Anthropic
2026-05-28
Claude Fable 5
Anthropic
2026-06-09
Claude Sonnet 5
Anthropic
2026-06-30
Claude Opus 5
Anthropic
2026-07-24
Claude Fable 5.1
Anthropic
2026-09-01
Claude Mythos 5.1
Anthropic
2026-09-01
Meta
Muse Spark
Meta
2026-04-08
Muse Image
Meta
2026-07-07
Muse Spark 1.1
Meta
2026-07-09
Muse Spark 1.2
Meta
2026-08-05
Muse Code
Meta
2026-08-05
Muse Glimmer
Meta
2026-08-10
30B
Muse Spark 1.3
Meta
2026-09-02
Microsoft
MAI-Code-1-Flash
Microsoft
2026-06-02
MAI-Image-2.5
Microsoft
2026-06-02
MAI-Thinking-1
Microsoft
2026-06-02
Mistral
Mistral 7B
Mistral
2023-09-27
Mixtral 8x7B
Mistral
2023-12-11
12.9B active / 46.7B totalMoE
Mistral Large
Mistral
2024-02-26
Mistral Nemo
Mistral
2024-07-18
12B
Mistral Large 2
Mistral
2024-07-24
123B
Pixtral 12B
Mistral
2024-09-17
12B
Pixtral Large
Mistral
2024-11-18
124B
Mistral Small 3
Mistral
2025-01-30
24B
Magistral Medium
Mistral
2025-06-10
Mistral 3 Family
Mistral
2025-12-02
675B totalMoE
Mistral Small 4
Mistral
2026-03-16
6B active / 119B totalMoE
Mistral Medium 3.5
Mistral
2026-04-30
~128B
X.AI
Grok-1.5
X.AI
2024-03-29
Grok-2
X.AI
2024-08-14
Grok-3
X.AI
2025-02-17
Grok-3 Mini
X.AI
2025-02-17
Grok 4
X.AI
2025-07-09
Grok Code Fast
X.AI
2025-08-28
Grok 4.1
X.AI
2025-11-17
Grok Imagine 1.0
X.AI
2026-01-28
Grok 4.3
X.AI
2026-04-30
Grok 4.5
X.AI
2026-07-08
Grok 4.6
X.AI
2026-08-12
Deep Seek
DeepSeek-Coder
Deep Seek
2023-11-02
DeepSeek-MoE
Deep Seek
2024-01-01
2.8B active / 14.6B totalMoE
DeepSeek-V2
Deep Seek
2024-05-01
21B active / 236B totalMoE
DeepSeek-V3
Deep Seek
2024-12-25
37B active / 671B totalMoE
DeepSeek-R1
Deep Seek
2025-01-20
37B active / 671B totalMoE
DeepSeek-V3.1
Deep Seek
2025-08-21
37B active / 671B totalMoE
DeepSeek-Terminus
Deep Seek
2025-09-22
DeepSeek-V3.2
Deep Seek
2025-12-01
37B active / 671B totalMoE
DeepSeek-V4-Pro
Deep Seek
2026-04-24
49B active / 1.6T totalMoE
DeepSeek-V4-Flash
Deep Seek
2026-04-24
13B active / 284B totalMoE
DeepSeek-V4-Flash-0731
Deep Seek
2026-07-31
13B active / 284B totalMoE
DeepSeek-V4-Flash-Vision-Exp
Deep Seek
2026-08-21
13B active / 284B totalMoE
DeepSeek-V4-Pro-0813
Deep Seek
2026-08-13
49B active / 1.6T totalMoE
DeepSeek-V4.1-Flash
Deep Seek
2026-09-10
16B active / 552B totalMoE
Alibaba
Qwen-Image-3.0
Alibaba
2026-07-21
Qwen-7B
Alibaba
2023-08-03
7B
Qwen 1.5
Alibaba
2024-02-04
Qwen2
Alibaba
2024-06-07
72B
Qwen2.5
Alibaba
2024-09-19
72B
QwQ-32B-Preview
Alibaba
2024-11-27
32B
Qwen3
Alibaba
2025-04-29
22B active / 235B totalMoE
Qwen3-Coder
Alibaba
2025-07-15
Qwen3.5
Alibaba
2026-02-16
17B active / 397B totalMoE
Qwen3-Max-Thinking
Alibaba
2026-01-27
Happy Horse 1.0
Alibaba
2026-04-07
~15B
Qwen3.8-Max-Preview
Alibaba
2026-07-19
2.4T totalMoE
Qwen3.8-Max
Alibaba
2026-08-02
2.4T totalMoE
Qwen3.8-2.4T-A95B
Alibaba
2026-08-12
95B active / 2.4T totalMoE
Qwen3.7-Max
Alibaba
2026-05-20
Qwen3.6 Plus
Alibaba
2026-03-31
ByteDance
Doubao Seed 2.0 Pro
ByteDance
2026-02-14
Doubao 2.1 Pro
ByteDance
2026-06-23
Seedance 2.0
ByteDance
2026-02-12
Seedance 2.5
ByteDance
2026-06-23
Seedream 3.0
ByteDance
2025-04-16
Seedream 4.0
ByteDance
2025-09-09
Seedream 4.5
ByteDance
2025-12-04
Seedream 5.0 Pro
ByteDance
2026-07-08
Seedance 1.0
ByteDance
2025-06-11
Tencent
Hunyuan-T1
Tencent
2025-03-21
Hy3
Tencent
2026-07-06
21B active / 295B totalMoE
Hy4-preview
Tencent
2026-08-28
49B active / 770B totalMoE
Moonshot AI
Kimi K2.5
Moonshot AI
2026-01-27
32B active / 1T totalMoE
Kimi K3
Moonshot AI
2026-07-16
50B active / 2.8T totalMoE
Zhipu AI
GLM-5
Zhipu AI
2026-02-11
40B active / 744B totalMoE
GLM-5.2
Zhipu AI
2026-06-13
40B active / 753B totalMoE
GLM-5.3
Zhipu AI
2026-08-14
40B active / 753B totalMoE
Xiaomi
MiMo-V2.5-Pro
Xiaomi
2026-04-27
42B active / 1.0T totalMoE
Cohere
Command A+
Cohere
2026-05-20
25B active / 218B totalMoE
Command R+
Cohere
2024-04-24
104B
Aya Expanse 32B
Cohere
2024-10-24
32B
Command A
Cohere
2025-03-13
111B
Baidu
ERNIE Bot
Baidu
2023-03-16
ERNIE 4.0
Baidu
2023-10-17
ERNIE 4.0 Turbo
Baidu
2024-06-28
ERNIE 4.5
Baidu
2024-12-20
ERNIE X1
Baidu
2025-02-28
ERNIE 5.0
Baidu
2026-01-22
72B active / 2.4T totalMoE
ERNIE 5.1
Baidu
2026-05-08
36B active / 800B totalMoE
Kuaishou
Kling 1.0
Kuaishou
2024-06-06
Kling 1.5
Kuaishou
2024-09-19
Kling 1.6
Kuaishou
2024-12-19
Kling 2.0
Kuaishou
2025-04-15
Kling 2.5 Turbo
Kuaishou
2025-09-23
Kling 2.6
Kuaishou
2025-12-03
Kling 3.0
Kuaishou
2026-02-05
Midjourney
Midjourney V5
Midjourney
2023-03-15
Midjourney V6
Midjourney
2023-12-21
Midjourney V6.1
Midjourney
2024-07-30
Midjourney V7
Midjourney
2025-04-03
Midjourney V8 Alpha
Midjourney
2026-03-17
Black Forest Labs
Flux 1.1 Pro
Black Forest Labs
2024-10-02
Flux.1 Tools
Black Forest Labs
2024-11-21
Flux 2.0
Black Forest Labs
2025-11-25
FLUX 3
Black Forest Labs
2026-07-23
Reve
Reve 2.0
Reve
2026-06-03
Reve 2.1
Reve
2026-07-08
Runway
Runway Gen-3 Alpha
Runway
2024-06-18
Runway Gen-4
Runway
2025-04-01
Gen-4.5
Runway
2025-12-01
GWM-1
Runway
2025-12-11
World Labs
Marble
World Labs
2025-11-12
Atlas
World Labs
2026-09-01
Odyssey
Odyssey-2
Odyssey
2025-10-27
Physical Intelligence
π0.7
Physical Intelligence
2026-04-16
LLM throne — ChatGPT (2022-11-30)LLM throne — GPT-4 (2023-03-14)LLM throne — GPT-4 Turbo (2023-11-06)LLM throne — Gemini 1.5 Pro (2024-02-15)LLM throne — Claude 3.5 Sonnet (2024-06-21)LLM throne — OpenAI o1 (2024-09-12)LLM throne — Grok-3 (2025-02-17)LLM throne — Claude 3.7 Sonnet (2025-02-24)LLM throne — o3 (2025-04-16)LLM throne — Claude Opus 4 (2025-05-22)LLM throne — Gemini 3 (2025-11-18)LLM throne — Claude Opus 4.5 (2025-11-24)LLM throne — Claude Opus 4.6 (2026-02-05)LLM throne — Gemini 3.1 Pro (2026-02-19)LLM throne — GPT-5.4 Thinking (2026-03-05)LLM throne — Claude Mythos Preview (2026-04-07)LLM throne — Claude Fable 5 (2026-06-09)LLM throne — too close to call: Claude Fable 5.1 trails Claude Fable 5 by 0.31 of the 1-point band. The crown stays with the earlier model until a challenger clears it.
How are thrones calculated?

Up to 2025, the 👑 LLM throne went to the highest mean of four benchmarks (GPQA, HLE, SWE-Bench Verified, MMLU). From 2026 on a model enters the race with a coding signal — SWE-Bench Pro or Terminal-Bench 2.x — plus GPQA and at least three of the four current-era values (GPQA, SWE-Bench Pro, Terminal-Bench 2.x, MMLU Pro). Scores are min-max normalised across the race pool before they are averaged, so a model gains nothing by omitting the benchmark it would score worst on.

A challenger has to clear the incumbent by 1 normalised point to take the crown, and only a model released after the incumbent can challenge. Inside that band the crown stays put and the line shows a dashed stub to a hollow dot — a live near-tie, not a handover. The 🏆 Coding throne is the same construction over the two coding signals alone (Verified pre-2026).

The 🎬 Video and 🎨 Image thrones use community Elo from the LMArena Text-to-Video and Text-to-Image arenas — the highest-rated model we track wins (with a 95% CI tie-break), since no harmonised academic benchmark exists across video/image generators. Models that aren't on the arena (e.g. Midjourney) can't hold the throne.

LLM and Coding thrones are derived from our own registry — every benchmark score is entered and sourced by hand, with no automated sync. Image and Video thrones are synced from the LMArena leaderboard dataset. Full write-up on the methodology page; for cross-vendor comparisons see artificialanalysis.ai.

Latest model tracked: Sep 10, 2026

2026
Sep
Deep Seek
552B total·16B activeMoE

Natively multimodal successor to DeepSeek-V4-Flash-0731 and the first model of DeepSeek's Causal Encoder-Decoder family: 40 layers split into a 20-layer causal encoder and a 20-layer decoder, a 552B backbone plus 196B of sparsely-accessed Engram memory, activating 8B parameters per token during prefill and 16B during decode. 1M-token context, MIT license, open weights. FP4 main KV caching at 890 bytes/token and DSpark speculative decoding drive a reported 409.5 tokens/s end-to-end, and DeepSeek prices it at roughly a quarter of V4-Pro.

GPQA90.9%
HLE36.8%

Key Capabilities

  • 552B backbone MoE, 8B active (prefill) / 16B active (decode)
  • Native multimodal image and text input
  • 1M token context window
  • 384 routed experts (1 shared, 6 routed per token)
  • MIT license, open weights

Innovations

  • Causal Encoder-Decoder (CED) architecture
  • Compressed Sparse Attention 2 with hierarchical sparse indexer
  • 196B sparsely-accessed Engram memory
  • FP4 main KV cache at 890 bytes/token
  • DSpark speculative decoding
  • Single-Pass mHC

GPT-6 Astra

2026-09-03
OpenAI
Architecture not disclosed

OpenAI's GPT-6 generation flagship, launched 2026-09-03 in a phased rollout: participants in OpenAI's cybersecurity access program first, then ChatGPT Plus, Pro, Business and Enterprise, the API (model id gpt-6-astra), Azure and AWS Bedrock. 1.05M-token context window, 128K max output, knowledge cutoff 2026-04-30, reasoning effort selectable from none to max. First OpenAI model classified 'Critical' for cybersecurity under the Preparedness Framework: the general release ships with offensive-cyber behaviour constrained, and a fuller cyber variant is limited to vetted defenders. Priced at $10/M input and $50/M output ($1.00 cached input, $12.50 cache write); requests above 272K input tokens are billed at 2x input and 1.5x output for the whole request, and a Fast mode costs 2x. OpenAI's launch table reports Terminal-Bench 4.0, Terminal-Bench-Science, OSWorld 2.0, FrontierMath Tier 4, DeepSWE v1.1 and HLE with tools, but no SWE-bench Pro, Terminal-Bench 2.x, MMLU-Pro or no-tools HLE figure, so the tracked fields here come from independent runs (Artificial Analysis, Vals AI, ARC Prize) where one exists.

GPQA96.1%
ARC-AGI-197.5%
ARC-AGI-295%

Key Capabilities

  • Frontier reasoning
  • Agentic coding
  • Computer use
  • Browser use
  • Tool use / function calling
  • Cybersecurity tasks (vulnerability discovery, exploit development)
  • Vision input

Innovations

  • API: gpt-6-astra
  • 1.05M-token context window
  • Preparedness Framework 'Critical' cyber classification with a gated cyber variant
  • Reasoning effort none to max, plus Fast mode at 2x price
  • Cache reads at 10% of input price
Google
Architecture not disclosed

Third Flash-tier release in six weeks and, per the model card, based on Gemini 3.7 Flash rather than a new training run. Google calls it its most intelligent Flash model yet, with the gains concentrated in software engineering and agentic knowledge work, and describes the model as spending extra reasoning steps and iterative tool calls on hard tasks instead of answering early. 1M-token context, 64K output, knowledge cutoff March 2026, three effort levels. Live at launch in the Gemini app, AI Mode, AI Studio, the Gemini API, Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported (model card, September 2026): DeepSWE v1.1 73.7, Terminal-bench 2.1 89.4, Terminal-bench 4.0 19.1 (a newer generation, tracked on its own axis), HLE-Verified 54.9, Vals Finance Agent v2 61.4, Harvey's Legal Agent Benchmark 10.0, OSWorld-2.0 59.0, CharXiv Reasoning 86.2, LABBench2 86.2, GDPval-AA v2 Elo 1545; Google's API documentation adds SWE-Bench Pro 61.6 and SWE-Atlas 51.9. Introductory API pricing of $0.75 / $3.75 per MTok until 2026-12-31, then $1.50 / $7.50. Announced together with Gemini 3.8 Flash Cyber.

GPQA94.44%
HLE47.82%
SWE-bench Pro61.6%
MMLU Pro90.22%

Key Capabilities

  • Multimodal input (text, image, audio, video)
  • Agentic workflows and coding
  • 1M-token context, 64K output
  • Configurable effort levels (low / medium / high)
  • Available via Gemini API, AI Studio, Antigravity and the Gemini Enterprise Agent Platform

Innovations

  • Iterates on Gemini 3.7 Flash three weeks after its release (model card: 'based on Gemini 3.7 Flash')
  • Extra reasoning steps and iterative tool calls on complex tasks, at the cost of more output tokens at higher effort levels
Google
Architecture not disclosed

Cybersecurity-specialised variant of Gemini 3.8 Flash and successor to Gemini 3.5 Flash Cyber, built to discover, validate and patch software vulnerabilities autonomously across codebases in 20 programming languages. Not a public model: access runs through Google's new Fairwind Program, which pairs the model with the CodeMender harness for government agencies and national cyber authorities, critical-infrastructure operators (healthcare, telecom, energy, finance), maintainers of widely used software platforms and vetted security partners, under operational requirements such as multi-factor authentication and access limited to internal security staff. Google says it will not be generally released. Vendor-reported, none of it a tracked academic benchmark: CyberGym (autonomous vulnerability discovery) above 3.5 Flash Cyber and larger frontier models, no figure published; a success rate above 70% on an internal 20-language vulnerability-discovery benchmark; CWE-Bench (Collinear) patching at 47.2% pass@1; and 2.6x more correct Chrome patches than the best commercial models in the Chrome Security team's own evaluation. No standard academic benchmarks were reported.

Key Capabilities

  • Vulnerability discovery, validation and patching across 20 programming languages
  • Runs inside the CodeMender harness
  • Restricted access: Fairwind Program only (governments, critical infrastructure, core software maintainers, vetted security partners)

Innovations

  • Cybersecurity-specialised variant of Gemini 3.8 Flash; successor to Gemini 3.5 Flash Cyber
  • First model distributed through Google's Fairwind Program, a gated access scheme with operational requirements instead of a general release

Muse Spark 1.3

2026-09-02
Meta
Architecture not disclosed

Fourth Muse Spark release in five months, live at launch in Muse Code and the Meta Model API. Meta frames the update as an efficiency change as much as a capability one: for the same agentic work it reports about 20% fewer tool calls and about 25% fewer tokens, with fewer turns and less verbose output. 1M-token context, multimodal input; xhigh reasoning effort is the generally available tier, a max tier is in limited preview for Meta partners. Vendor-reported (Meta launch scorecard, September 2026): Terminal-Bench 2.1 88.8, DeepSWE v1.1 75.4, SWEAtlas CodeBase QnA 59.4, MRCR 256K-512K 98.5 and MRCR 512K-1M 98.1. API pricing unchanged from 1.1 and 1.2 at $1.25 / $4.25 per million input/output tokens with cached input at $0.15; the contributor tier (over 90% discount in exchange for training on prompts and completions) continues.

GPQA94%
HLE47%

Key Capabilities

  • Agentic coding
  • Long-horizon tool use
  • Multimodal input
  • 1M-token context
  • Reasoning effort control (xhigh; max in limited preview)

Innovations

  • Token and tool-call efficiency as a stated release goal, at unchanged price and context
  • Fourth frontier release from Meta Superintelligence Labs in five months
Anthropic
Architecture not disclosed

Successor to Claude Fable 5 in the Mythos-class tier, at the same $10/M input and $50/M output pricing but with cache reads cut to $0.25/M (2.5% of the input price instead of 10%). Launched alongside Claude Mythos 5.1, the same underlying model with safeguards lifted for Project Glasswing participants (tracked separately). Safety classifiers remain in place; refused requests can fall back server-side to Claude Opus 4.8 or Claude Opus 5. Adaptive thinking is always on, forced tool use is no longer supported, and all text output carries Anthropic's statistical watermark. Vendor-reported benchmarks: Terminal-Bench 4.0 55.8 (Fable 5: 42.0, Opus 5: 52.3), Terminal-Bench-Science 0.1 52.6, HLE 60.9 without tools (65.0 with tools), OSWorld 2.0 41.7 strict, CursorBench 3.2.0 73.4, AutomationBench 31.4.

GPQA93.43%
HLE60.9%
MMLU Pro92.38%

Key Capabilities

  • Autonomous long-horizon tasks
  • Agentic coding
  • Multistep research
  • Computer use
  • Vision input
  • Tool use

Innovations

  • API: claude-fable-5-1
  • Cache reads at 2.5% of input price
  • Content provenance (text watermark, C2PA for files)
  • Per-message effort control
Anthropic
Architecture not disclosed

Same underlying model as Claude Fable 5.1, offered by invitation only to Project Glasswing participants and the Cyber and Life Sciences verification programs, with the safety classifiers and fallback routing that govern Fable 5.1 lifted. Successor to Claude Mythos 5. Shares Fable 5.1's specifications and pricing ($10/M input, $50/M output, 1M context). API ID claude-mythos-5-1 on the Claude API, Bedrock, Google Cloud and Foundry; no public access. Anthropic reports Terminal-Bench 4.0 60.9 against 55.8 for Fable 5.1 — the gap is the cost of the safeguard interventions — and HLE 60.9 without tools (65.0 with tools) for both models. Independent evaluators measure Fable 5.1, not this variant, so no independent values are carried over.

HLE60.9%

Key Capabilities

  • Autonomous long-horizon tasks
  • Agentic coding
  • Cyber defense
  • Multistep research
  • Computer use
  • Vision input

Innovations

  • API: claude-mythos-5-1
  • Project Glasswing restricted access
  • Safeguards lifted for verified defenders

Atlas

2026-09-01
World Labs
Architecture not disclosed

World Labs' omni world model, pretrained from scratch to operate natively on text, images, video and 3D as a multimodal autoregressive diffusion transformer with every input in one shared spatial context. The headline capability is camera-controlled video, up to one minute at 1440p from one or more reference images, with the camera path supplied as a geometric input rather than described in a prompt. The same model reconstructs scenes into point clouds and 3D Gaussian splats, works with depth maps and camera poses, reframes multi-camera footage and supports parts of a real-to-sim robotics workflow. Early access for selected partners via a request form; no pricing, API or public date announced. World Labs reports that third-party human raters preferred Atlas over rival video models in 75 to 94 percent of pairwise trials depending on the competitor, a vendor-commissioned preference study recorded here as a note, not as a score.

Key Capabilities

  • Camera-controlled video up to one minute at 1440p
  • Text, image, video and 3D inputs in one spatial context
  • 3D reconstruction to point clouds and Gaussian splats
  • Real-to-sim workflows for robotics

Innovations

  • Single omni model for generation, reconstruction and simulation
  • Camera path as geometric input instead of prompt text
Aug

Hy4-preview

2026-08-28
Tencent
770B total·49B activeMoE

Preview of Tencent's next-generation Hunyuan flagship: a 770B-parameter MoE activating 49B per token, with a 1M-token context window. Open-weight under Apache 2.0 with an FP8 checkpoint alongside BF16, tuned for coding agents, complex tool-use workflows and productivity tasks, and trained on domain data from Tencent's software, gaming and finance teams.

GPQA92.3%
SWE-bench Pro65.7%

Key Capabilities

  • 770B total / 49B active MoE
  • 1M token context window
  • Coding-agent and tool-use focus
  • Apache 2.0 license

Innovations

  • Tencent-internal software, gaming and finance domain data
  • FP8 checkpoint shipped alongside BF16
Google
Architecture not disclosed

Successor to Gemini Omni Flash (2026-06-30) for video generation and editing, released as the model id gemini-omni-1.1-flash. Takes any combination of text, image, audio and video as input and returns video. Scene extension now conditions on up to 10 seconds of preceding footage instead of the tail of the clip and continues in 10-second increments up to 40 seconds total; a first-and-last-frame control renders a single continuous shot between two supplied frames, and a video reference of up to 3 seconds carries character and style across generations. Native output resolution is 720p, with 1080p and 4K delivered as upscales, plus a new 360p draft mode that Google reports as up to 60% faster at a third of the 720p cost. Per-second API pricing: $0.03 (360p), $0.10 (720p), $0.15 (1080p), $0.30 (4K). Live via the Gemini API, Google AI Studio, Flow and the Gemini Enterprise Agent Platform; scene extension is rolling out to Google AI Plus, Pro and Ultra subscribers in the Gemini app globally. The predecessor's gemini-omni-flash-preview endpoint is scheduled for deprecation on 2026-09-30.

Key Capabilities

  • Text-to-video
  • Image/audio/video-to-video (any-to-any input)
  • Conversational multi-turn editing
  • Scene extension up to 40s in 10s increments
  • First-and-last-frame shot control
  • Video reference up to 3s for character/style consistency
  • 720p native output, 1080p and 4K upscale
  • 360p draft mode

Innovations

  • Scene extension conditioned on up to 10s of prior footage, extending to 40s total
  • First-and-last-frame control for a single continuous shot
  • 360p draft mode at a third of the 720p per-second price
  • 4K upscaled export at $0.30 per second
Deep Seek
284B total·13B activeMoE

Experimental vision variant of DeepSeek-V4-Flash-0731, live on the DeepSeek API platform as model 'deepseek-v4-flash-vision-exp'. Same 284B total / 13B active MoE and 1M-token context as the text build, extended to image input (text output only); images are billed as up to 384 tokens each and can be passed as base64 through Chat Completions, Messages and the Responses API. DeepSeek describes text capability — agents, reasoning, world knowledge — as unchanged versus V4-Flash and reports the gains on agent benchmarks that require visual understanding: Terminal-Bench 2.1 83.9 and DeepSWE 59.3. Unlike the rest of the V4 line this is a closed API release — no open weights, no Hugging Face repository at launch, and the 'Exp' tag marks it as experimental.

Key Capabilities

  • 284B total / 13B active MoE
  • Image input, text output
  • 1M token context window
  • Multimodal agent tasks
  • Chat Completions, Messages and Responses API
  • Images billed at up to 384 tokens each

Innovations

  • First vision-enabled model of the DeepSeek V4 line
  • Vision added on top of the V4-Flash-0731 checkpoint at unchanged parameter count
  • API-only release — first V4 build without open weights

GLM-5.3

2026-08-14
Zhipu AI
753B total·40B activeMoE

Post-training-only update to GLM-5.2: Z.ai leaves the 753B MoE base model unchanged and attributes the capability gains entirely to scaled post-training. Weights landed on Hugging Face on 2026-08-28, fourteen days after launch, once Z.ai's safety review completed — but under a bespoke GLM-5.3 License instead of the MIT license GLM-5.2 shipped under: use, modification, distribution, sublicensing, sale, deployment and fine-tuning are all permitted, but a company with more than $10B aggregate revenue over any 12 consecutive months must pass a Z.ai security review before hosting the model commercially. Individual users and smaller companies are unaffected. Z.ai reports its results on a new generation of agentic benchmarks rather than the axes tracked here (maximum thinking effort): Terminal-Bench 3.0 28.3 (GLM-5.2: 4.6), DeepSWE v1.1 66.9 (46.2), Agents' Last Exam 28.5 (23.8), CyberGym 84.5, plus AutomationBench, GDPVal-AA v2 and HLE with tools. Only Terminal-Bench 3.0 has a column here (its own axis, separate from the 2.x scale); the rest do not map onto the tracked benchmarks, and Z.ai published no GPQA, MMLU-Pro or SWE-bench Pro table.

Key Capabilities

  • 753B total / ~40B active MoE (base unchanged from GLM-5.2)
  • 1M-token context
  • Agentic / long-horizon coding
  • Cybersecurity workflows
  • Open weights under the custom GLM-5.3 License (revenue-tiered)

Innovations

  • Capability gain from post-training scaling alone, base model unchanged
  • Revenue-gated open-weight license replacing MIT (security review above $10B revenue)
Deep Seek
1.6T total·49B activeMoE

General-availability build of DeepSeek-V4-Pro, replacing the April preview. Same 1.6T total / 49B active MoE architecture as the preview — re-post-trained, with the gains concentrated in agentic and software-engineering tasks. The deepseek-v4-pro API endpoint now points to this build. 1M-token context, MIT license.

GPQA92.83%
HLE39.34%
SWE-bench96.4%
MMLU Pro86.97%

Key Capabilities

  • 1.6T total / 49B active MoE
  • 1M token context window
  • Three reasoning effort modes (incl. Think Max)
  • MIT license

Innovations

  • Hybrid attention: CSA + HCA
  • Manifold-Constrained Hyper-Connections
  • Re-post-training of the April preview checkpoint
Google
Architecture not disclosed

Workhorse successor to Gemini 3.6 Flash, shipped three weeks after it. Google states the model was not trained from scratch — it replaces the predecessor via algorithmic improvements and user feedback. Same 1,048,576-token context window as 3.6 Flash, multimodal input (text, image, audio, video, files). Live through the Gemini API in AI Studio, Android Studio, Google Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported coding deltas against 3.6 Flash: DeepSWE v1.1 49.0 to 65.3, FrontierCode 1.1 Main 34.4 to 43.6. Introductory API pricing of $0.75 / $3.75 per MTok runs until 2026-12-31, after which Google lists $1.50 / $7.50.

GPQA93.94%
HLE47.87%
SWE-bench Pro60.4%
MMLU Pro90.12%

Key Capabilities

  • Multimodal input (text, image, audio, video, files)
  • Agentic workflows and coding
  • 1M-token context
  • Available via Gemini API, Android Studio and Antigravity

Innovations

  • Successor built by replacing 3.6 Flash through algorithmic improvements rather than a from-scratch training run
Alibaba
2.4T total·95B activeMoE

Open-weight checkpoint of the Max-class flagship and the first Qwen-Max-class model with downloadable weights: 2.4T total / 95B active sparse MoE, published on Hugging Face and ModelScope alongside an FP8 build and served via API. It is not the hosted Qwen3.8-Max — the checkpoint is text-only (no vision or video input) and always reasons, thinking mode cannot be disabled. Native context is 262,144 tokens, extensible to roughly 1M via YaRN. Shipped under the bespoke Qwen3.8-Max License rather than Apache 2.0. In FP8 the weights occupy about 2,325 GiB, so self-hosting takes a 16-GPU class deployment.

GPQA93.54%
HLE42.45%

Key Capabilities

  • 2.4T total / 95B active parameters (sparse MoE)
  • 262K native context, ~1M via YaRN
  • Text-only input and output
  • Always-on thinking (cannot be disabled)
  • 128K max output tokens
  • FP8 and BF16 checkpoints, vLLM / SGLang support

Innovations

  • First Qwen-Max-class model released with open weights
  • 2.4T-parameter open-weight MoE
  • Fine-grained FP8 block quantization (block size 128)

Grok 4.6

2026-08-12
X.AI
Architecture not disclosed

Reasoning-first flagship succeeding Grok 4.5, with a 500K-token context window and text + image input. Same $2 / $6 per million input/output tokens as Grok 4.5, with a 75% cache-hit discount.

GPQA94.7%
HLE42.9%
SWE-bench95.6%
MMLU Pro89.4%

Key Capabilities

  • 500K token context window
  • Agentic coding
  • Reasoning-first
  • Text and image input

Innovations

    Muse Glimmer

    2026-08-10
    Meta
    30B

    30B dense agentic model distilled from Muse Spark 1.2 and released under Apache 2.0 — Meta's first open-weight release since the Llama 4 family. Served through the API as well as downloadable: at 4-bit the checkpoint stays under 20 GB, so the full setup including KV cache and perception encoder fits a 24–32 GB envelope on a single consumer GPU or Mac. A separate perception encoder handles image input; block-level speculative decoding keeps latency inside a real agent loop. 131K-token context.

    GPQA83.5%
    HLE21.96%
    SWE-bench76%
    SWE-bench Pro51.2%

    Key Capabilities

    • 30B dense parameters
    • Multimodal input via separate perception encoder
    • Agentic tool calling with planning and failure recovery
    • Controllable reasoning effort levels
    • 131K-token context
    • Apache 2.0

    Innovations

    • Distilled from Muse Spark 1.2
    • Block-level speculative decoding for agent-loop latency
    • 4-bit K-Quant checkpoint under 20 GB
    • Meta's first open-weight release since Llama 4

    Muse Spark 1.2

    2026-08-05
    Meta
    Architecture not disclosed

    Coding-focused update of the Muse Spark frontier family, released alongside and co-trained with the Muse Code terminal agent. Third Meta Superintelligence Labs model in four months. Meta evaluates it inside its own Muse Code harness at xhigh reasoning effort. Same API pricing as Muse Spark 1.1 at $1.25 / $4.25 per million input/output tokens, with cached input at $0.15; a contributor tier cuts the rate by over 90% in exchange for permission to train on prompts and completions. Rate limits reach 3.000 requests and 4M tokens per minute per team.

    GPQA90.4%
    HLE43.9%

    Key Capabilities

    • Agentic coding
    • Long-horizon tool use
    • Multimodal input
    • Multi-agent orchestration
    • Reasoning effort control (xhigh)

    Innovations

    • Co-trained with the Muse Code agent harness
    • Contributor tier (data-sharing discount)

    Muse Code

    2026-08-05
    Meta
    Architecture not disclosed

    Meta’s first coding agent: a terminal-based harness for macOS and Linux, released in beta and powered by Muse Spark 1.2. Runs a simple agent loop plus persistent async background agents that stay alive for a whole session and accumulate context instead of restarting per task; larger tasks are split onto parallel subagents in isolated git worktrees so the working copy stays untouched. Meta reports one run optimising Nvidia Hopper GPU kernels across more than 1.000 tool calls in sessions of up to 24 hours. Pay-as-you-go at Muse Spark 1.1 API rates, or a contributor tier at over 90% discount in exchange for training on prompts and completions.

    Key Capabilities

    • Terminal coding agent
    • Repository-scale planning, editing and validation
    • Persistent async background agents
    • Parallel subagents in isolated worktrees
    • Long-running sessions (24h+)

    Innovations

    • Session-persistent background agents
    • Contributor tier (data-sharing discount)

    Qwen3.8-Max

    2026-08-02
    Alibaba
    2.4T totalMoE

    General-availability build of the flagship Max model, two weeks after the WAIC preview. 2.4T-parameter sparse Mixture-of-Experts, text/image/video in and text out, 1M-token context (max ~991K input, ~983K with thinking enabled; 65,536 output tokens). Served via Alibaba Cloud Model Studio with an OpenAI- and DashScope-compatible API at $2 / $6 per MTok flat across the full context. The active-parameter count is still undisclosed and no benchmark table has been published. Alibaba announced open weights for Qwen3.8-Max plus a smaller Qwen3.8-27B checkpoint for the following week — the first Qwen-Max-class model to go open-weight. Both have since shipped and are tracked as separate local entries: Qwen3.8-2.4T-A95B (2026-08-12, text-only and always-thinking, custom Qwen3.8-Max License) and Qwen3.8-27B (2026-08-14, Apache 2.0). The open checkpoint is not identical to this hosted API — it drops the vision path — so the benchmark values here, measured against the API, are not carried over to it.

    GPQA92.73%
    HLE41.4%
    MMLU Pro88.6%

    Key Capabilities

    • 2.4T total parameters (sparse MoE)
    • 1M token context window
    • Multimodal input (text, image, video), text output
    • 65,536 max output tokens
    • OpenAI- and DashScope-compatible API
    • Open weights released as separate checkpoints (2.4T-A95B, 27B)

    Innovations

    • 2.4T Sparse MoE
    • First open-weight release of a Qwen-Max-class model
    • Flat pricing across the full 1M-token context
    Jul
    Deep Seek
    284B total·13B activeMoE

    General-availability build of DeepSeek-V4-Flash, replacing the April preview. Same 284B total / 13B active MoE architecture and size as the preview — only re-post-trained, with the gains concentrated in agentic tool use. 1M-token context, MIT license. On 2026-08-21 DeepSeek added a vision variant on top of this checkpoint, DeepSeek-V4-Flash-Vision-Exp — API-only, without open weights. Superseded on 2026-09-10 by DeepSeek-V4.1-Flash, which moves to a Causal Encoder-Decoder architecture rather than re-post-training this checkpoint.

    GPQA91%
    HLE37%
    MMLU Pro86.21%

    Key Capabilities

    • 284B total / 13B active MoE
    • 1M token context window
    • Three reasoning effort modes (incl. Think Max)
    • Responses API support
    • MIT license

    Innovations

    • Hybrid attention: CSA + HCA
    • Manifold-Constrained Hyper-Connections
    • Agent-focused re-post-training on the preview checkpoint
    Google
    Architecture not disclosed

    Second generation of Google DeepMind's robotics stack, extending control from the upper body to the whole body: a humanoid can walk, crouch, lean and manipulate in one coordinated motion, and several robots can collaborate on a task. Ships as three models — Gemini Robotics-ER 2 (reasoning, available in Google AI Studio and in private preview on the Gemini Enterprise Agent Platform), plus the VLA and an on-device variant for early-access partners.

    Key Capabilities

    • Whole-body humanoid control
    • Advanced dexterity
    • Multi-robot collaboration
    • On-device variant

    Innovations

    • Whole-body control including legs, not just manipulation
    • Multi-robot task coordination

    Claude Opus 5

    2026-07-24
    Anthropic
    Architecture not disclosed

    Opus-tier Claude flagship for agentic coding, long-horizon tool use and computer use, released as the new top model in the Claude apps and the API. Priced at $5 / $25 per million input/output tokens.

    GPQA93.43%
    HLE52.6%
    SWE-bench96%
    SWE-bench Pro79.2%
    MMLU Pro91.59%

    Key Capabilities

    • Agentic coding
    • Computer use
    • Extended thinking
    • Tool use
    • Vision input

    Innovations

    • API: claude-opus-5
    • Effort control

    FLUX 3

    2026-07-23
    Black Forest Labs
    Architecture not disclosed

    Unified multimodal flow model that generates image, video and audio from a single backbone and can be extended to predict robot actions. Rollout is staged: FLUX 3 Video with native synchronised audio (clips up to 20 seconds, aspect ratios from 9:16 to 21:9) and the action model are in application-based early access, image generation is announced for the following weeks, and an open-weight FLUX 3 Dev backbone for later in the year. No public API tier, pricing, parameter count or benchmark methodology published at launch. FLUX 3 Video now ranks #2 in the Text-to-Video Arena at ~1,496 Elo, behind Gemini Omni Flash (~1,512) and ahead of Dreamina Seedance 2.0 (~1,478) — the first independent quality signal for the model.

    Key Capabilities

    • Text-to-video (up to 20s)
    • Native synchronised audio
    • Image-to-video
    • Video-to-video from reference clip
    • Keyframe-to-video transitions
    • Robot action prediction

    Innovations

    • Single flow backbone for image, video, audio and action
    • Video and native audio generated jointly
    • Text-to-Video Arena #2 at intake (~1,496 Elo)

    Qwen-Image-3.0

    2026-07-21
    Alibaba
    Architecture not disclosed

    Third-generation text-to-image model from Alibaba's Qwen team, focused on dense text and layout rendering. Accepts prompts up to 4,500 tokens, renders in-image text in 12 languages and 20+ fonts down to ~10px, and produces complex multi-element layouts (newspapers, storyboards, infographics, UI mockups, knowledge graphs) in a single pass, generating up to 9 images at once. Unlike the earlier Apache 2.0 releases (Qwen-Image 1.0, Aug 2025, and 2.0, each with a technical report), 3.0 shipped without open weights, benchmarks or a technical report. Available via chat.qwen.ai, Qwen Studio and Alibaba's API platform; API pricing not disclosed at launch.

    Key Capabilities

    • Text-to-Image
    • Prompts up to 4,500 tokens
    • In-image text rendering (12 languages, 20+ fonts, ~10px)
    • Single-pass complex layouts
    • Up to 9 images per generation

    Innovations

    • Long-prompt layout generation (4.5K tokens)
    • High-density multilingual text rendering
    • Closed release (no open weights, unlike prior Qwen-Image versions)
    Google
    Architecture not disclosed

    Workhorse model in the Gemini 3 series, based on Gemini 3.5 Flash, delivering better coding, knowledge work and multimodal performance at improved token-efficiency (Google reports ~17% fewer output tokens than 3.5 Flash). 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch across the Gemini app, AI Studio and Gemini API. Vendor-reported (model card): SWE-Bench Pro 58.7, Terminal-Bench 2.1 78.0, OSWorld-Verified 83.0, MLE-Bench 63.9, CharXiv Reasoning (with tools) 89.4, GDM-MRCR v2 (128k) 91.8. Google's launch benchmark table also lists GPQA Diamond 90.4 (unchanged vs 3.5 Flash). Announced alongside 3.5 Flash-Lite and 3.5 Flash Cyber.

    GPQA93.43%
    HLE38.3%
    SWE-bench Pro58.7%
    MMLU Pro89.28%

    Key Capabilities

    • Multimodal input (text, image, audio, video)
    • Agentic workflows and coding
    • 1M-token context
    • GA at launch (2026-07-21)

    Innovations

    • ~17% fewer output tokens than Gemini 3.5 Flash at comparable or better quality (token-efficiency)
    Google
    Architecture not disclosed

    Cost-efficient, low-latency tier of the Gemini 3 series, based on Gemini 3.1 Flash-Lite and optimized for high-volume, latency-sensitive tasks (translation, classification, document processing) as well as agentic workflows. 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch and rolling out to Google Search. Vendor-reported (model card): SWE-Bench Pro 54.2, Terminal-Bench 2.1 54.0, OSWorld-Verified 74.0, MLE-Bench 39.2, CharXiv Reasoning (with tools) 76.5, GDM-MRCR v2 (128k) 72.2. Reported output speed ~350 tokens/s.

    GPQA83.84%
    HLE17.5%
    SWE-bench Pro54.2%
    MMLU Pro85.84%

    Key Capabilities

    • Multimodal input (text, image, audio, video)
    • High-throughput, low-latency tasks
    • Agentic workflows
    • 1M-token context
    • Rolling out to Google Search

    Innovations

    • Output speed ~350 tok/s at $0.30 / $2.50 per 1M tokens (cost-efficient high-volume tier)
    Google
    Architecture not disclosed

    Security-specialized model fine-tuned from Gemini 3.5 Flash to identify, validate and patch software vulnerabilities at low cost per token. Deployed inside Google's CodeMender security agent, which calls it many times in rapid succession to explore alternative execution paths across a codebase. Availability is limited: an initial pilot for governments and trusted partners via CodeMender, not general API access. Google reports internal results of 55 confirmed unique vulnerabilities on Chrome's V8 JavaScript engine (vs 47 for standard 3.5 Flash and 36 for Claude Opus 4.6), a 42% improvement on long-range multi-turn cyber benchmarks over the Flash 3 predecessor, and frontier-competitive performance on CyberGym within CodeMender. No standard academic benchmarks were reported.

    Key Capabilities

    • Vulnerability detection, validation and patching
    • Runs inside the CodeMender security agent
    • Limited pilot: governments and trusted partners only

    Innovations

    • Cybersecurity-specialized fine-tune of Gemini 3.5 Flash for automated vulnerability discovery and patching
    Alibaba
    2.4T totalMoE

    Preview of Alibaba's flagship Max model, unveiled at WAIC Shanghai. 2.4T-parameter sparse Mixture-of-Experts, natively multimodal (text, images, video, documents) with a 1M-token context window (inherited from Qwen3.7-Max). Active-parameter count, full benchmark suite and license were not disclosed at preview. Runs at 10% of standard pricing during the preview via Token Plan / Qoder; an open-weight release was announced as forthcoming.

    Key Capabilities

    • 2.4T total parameters (sparse MoE)
    • 1M context window
    • Natively multimodal (text, image, video, docs)
    • Preview via Token Plan / Qoder
    • Open weights planned

    Innovations

    • 2.4T Sparse MoE
    • Multimodal Text/Image/Video/Docs
    • Open-Weight Max Release Planned

    Kimi K3

    2026-07-16
    Moonshot AI
    2.8T total·50B activeMoE

    2.8T MoE (~50B active, 16 of 896 experts) native multimodal flagship with a 1M-token context window. Launched via API, app and playground on 2026-07-16; open weights scheduled for 2026-07-27.

    GPQA93.5%
    HLE43.5%
    MMLU Pro87.97%

    Key Capabilities

    • 2.8T total / ~50B active MoE (896 experts, 16 routed)
    • 1M token context window
    • Natively multimodal (text, image, video)
    • Vision-in-the-loop (screenshot inspect + code edit)
    • Open weights scheduled 2026-07-27

    Innovations

    • Kimi Delta Attention
    • 2.8T Open-Weight MoE (largest open-weight at launch)
    • Vision-in-the-Loop Agent Feedback
    • 1M-Token Native Context

    Muse Spark 1.1

    2026-07-09
    Meta
    Architecture not disclosed

    Second Meta Superintelligence Labs model and the launch vehicle for the paid Meta Model API. Multimodal reasoning model built for agentic work: takes text, image, video, audio and PDF input, returns text, with a 1M-token context window. Runs as main agent that plans and delegates or as a subagent, and generalises zero-shot to new tools, MCP servers and custom skills. Free for consumers in the Meta AI app (Thinking mode); developer preview US-only at launch. $1.25 / $4.25 per million input/output tokens.

    GPQA91.2%
    HLE45.1%
    SWE-bench Pro61.5%
    MMLU Pro88.7%

    Key Capabilities

    • Multimodal input (text, image, video, audio, PDF)
    • 1M token context window
    • Agentic tool use
    • Multi-agent orchestration (main agent / subagent)
    • MCP server support
    • Reasoning

    Innovations

    • Meta Model API (first paid Meta inference API)
    • Zero-shot generalisation to unseen tools, MCP servers and skills

    Grok 4.5

    2026-07-08
    X.AI
    Architecture not disclosed

    Coding- and agentic-focused flagship from xAI, trained alongside Cursor. Available in Grok Build, Cursor and the console at $2 / $6 per million input/output tokens.

    GPQA92.93%
    HLE40.3%
    SWE-bench Pro64.7%
    MMLU Pro89.22%

    Key Capabilities

    • Agentic coding
    • Reasoning-first

    Innovations

      Reve 2.1

      2026-07-08
      Reve
      Architecture not disclosed

      Iteration on Reve's layout-first 4K text-to-image model, with improved prompt understanding, world knowledge and in-image text rendering (including foreign scripts). Retains the layout-as-prompt architecture and native 4K output.

      Key Capabilities

      • Layout-first generation
      • Native 4K (4096×4096) output
      • Addressable / editable elements
      • Text rendering in images
      • Foreign-script text rendering

      Innovations

      • Layout-as-prompt (position + size + local description)
      • Native 4096×4096 output
      ByteDance
      Architecture not disclosed

      Pro tier of ByteDance's Seedream 5.0 image family (above Seedream 5.0 and 5.0 Lite). Adds reasoning- and online-search-guided generation, separable-layer editing, and text rendering in 10+ languages, aimed at high-density infographics and professional editing.

      Key Capabilities

      • Text-to-Image
      • Reasoning-guided generation
      • Online-search-guided generation
      • Separable-layer editing
      • Text rendering in 10+ languages

      Innovations

      • Reasoning + online search for image generation
      • Layer-separable editing
      • Multilingual text rendering (10+ languages)

      Muse Image

      2026-07-07
      Meta
      Architecture not disclosed

      Meta Superintelligence Labs' first image-generation model (codename Mango). Operates agentically, invoking search and coding tools, self-refining its outputs and scaling test-time compute, for instruction following, precise editing and multi-reference composition. Available in the Meta AI app, on meta.ai, in Instagram Stories (US) and WhatsApp (limited).

      Key Capabilities

      • Text-to-Image
      • Precise image editing
      • Multi-reference composition
      • Agentic tool use (search + code)
      • Test-time self-refinement

      Innovations

      • First Meta Superintelligence Labs image model
      • Agentic generation (tool use + self-refinement)

      Hy3

      2026-07-06
      Tencent
      295B total·21B activeMoE

      Tencent's Hunyuan 3 flagship: a 295B-parameter MoE activating 21B per token (192 routed experts with top-8 routing plus an always-active shared expert), 256K context window and three selectable reasoning-effort modes. Full release under Apache 2.0 — unlike the April preview license, without territorial restrictions.

      GPQA90.4%
      SWE-bench78%
      SWE-bench Pro57.9%

      Key Capabilities

      • 295B total / 21B active MoE
      • 256K token context window
      • Three reasoning-effort modes
      • Apache 2.0 license

      Innovations

      • Dense-MoE hybrid with always-active shared expert
      • Multi-Token-Prediction layers (3.8B)
      Jun

      Claude Sonnet 5

      2026-06-30
      Anthropic
      Architecture not disclosed

      Sonnet-tier Claude model for agentic coding, computer use and tool use, released as the new default in the Claude apps and the API.

      GPQA88.89%
      HLE43.2%
      SWE-bench Pro63.2%
      MMLU Pro87.55%

      Key Capabilities

      • Agentic coding
      • Computer use
      • Extended thinking
      • Tool use
      • Vision input

      Innovations

      • API: claude-sonnet-5
      • OSWorld-Verified computer use
      • Introductory pricing through 2026-08-31
      Google
      Architecture not disclosed

      Faster, lower-cost variant of Nano Banana 2 for rapid text-to-image generation (~4s per image).

      Key Capabilities

      • Fast text-to-image (~4s)
      • Text rendering in images
      • Low per-image cost
      • Image editing

      Innovations

      • Lite tier of Gemini 3.1 Flash Image
      • ~$0.034 per generated image
      • Latency-optimized generation
      Google
      Architecture not disclosed

      Google's cost-efficient multimodal video-generation model with conversational editing; generates and refines video from text, image, audio and video inputs. 720p, up to 10s clips, $0.10/sec via the Gemini API.

      Key Capabilities

      • Text-to-video
      • Image/audio/video-to-video
      • Conversational multi-turn editing
      • 720p output
      • Up to 10s clips

      Innovations

      • Gemini Omni model family
      • Native conversational video editing
      • $0.10 per second (Gemini API / AI Studio)

      GPT-5.6 Sol

      2026-06-26
      OpenAI
      Architecture not disclosed

      Preview of OpenAI's GPT-5.6 frontier model (Sol tier), part of a Sol/Terra/Luna family in limited preview to trusted partners via Codex and the API. Sol priced at $5/M input, $30/M output tokens. Not yet generally available.

      GPQA95.2%
      HLE47.2%
      SWE-bench Pro64.6%
      MMLU Pro89.1%

      Key Capabilities

      • Frontier reasoning
      • Agentic coding
      • Tool use / function calling
      • Computer use
      • Cybersecurity tasks

      Innovations

      • Sol / Terra / Luna model family
      • Limited preview (Codex + API)
      • Token-efficient agentic workflows

      GPT-5.6 Terra

      2026-06-26
      OpenAI
      Architecture not disclosed

      Balanced mid-tier of OpenAI's GPT-5.6 Sol/Terra/Luna family, positioned as the everyday default for interactive and agentic coding. Priced at $2.50/M input, $15/M output tokens. Part of the limited preview from 2026-06-26; generally available 2026-07-09.

      Key Capabilities

      • Agentic coding
      • Tool use / function calling
      • Programmatic tool calling (Responses API)

      Innovations

      • Sol / Terra / Luna model family

      GPT-5.6 Luna

      2026-06-26
      OpenAI
      Architecture not disclosed

      Lightweight, cost-efficient tier of OpenAI's GPT-5.6 Sol/Terra/Luna family, aimed at smaller and faster tasks. Lowest-cost option in the family at $1/M input, $6/M output tokens. Part of the limited preview from 2026-06-26; generally available 2026-07-09.

      Key Capabilities

      • Agentic coding
      • Tool use / function calling
      • Programmatic tool calling (Responses API)

      Innovations

      • Sol / Terra / Luna model family

      Doubao 2.1 Pro

      2026-06-23
      ByteDance

      Auf der Volcano Engine Force Conference 2026 vorgestelltes Flaggschiff (offiziell Doubao-Seed-2.1 Pro), positioniert am production-level capability threshold. Schwerpunkte: Code-Delivery, Long-Horizon-Agent-Tasks, multimodales Verständnis und Enterprise-Stabilität. Proprietär über Volcano Engine. Unabhängige Benchmark-Werte stehen zum Eintragungszeitpunkt noch aus.

      Key Capabilities

      • Coding / Engineering Delivery
      • Long-Horizon Agent
      • Vision Language / Multimodal
      • Enterprise-Stabilität

      Innovations

      • Production-Threshold-Flaggschiff

      Seedance 2.5

      2026-06-23
      ByteDance
      Architecture not disclosed

      Beta preview of ByteDance's Seedance 2.5 video model, headline feature 30-second one-shot generation. Announced as a global enterprise beta at Volcano Engine 2026; public launch targeted for early July 2026.

      Key Capabilities

      • Text-to-video
      • 30s one-shot generation
      • Audio-video joint generation
      • Multimodal inputs

      Innovations

      • 30-second single-shot video
      • Enterprise beta (Volcano Engine)
      • Seedance 2.x architecture

      GLM-5.2

      2026-06-13
      Zhipu AI
      753B total·40B activeMoE

      753B Mixture-of-Experts (~40B active) open-weight model under MIT license. Introduces the IndexShare attention mechanism (~2.9x lower per-token FLOPs at 1M-token context) and a usable 1M-token context window. Ships with two thinking-effort levels; the highest tier is marketed as 'Max'. Available as open weights on Hugging Face and via the Z.ai API. Vendor-reported benchmarks (maximum thinking effort): GPQA Diamond 91.2, SWE-bench Pro 62.1, Terminal-Bench 2.1 81.0, HLE 40.5 (no tools).

      GPQA85.61%
      HLE40.5%
      SWE-bench Pro62.1%
      MMLU Pro86.71%
      Agent Arena+4.4%#10

      Key Capabilities

      • 753B total / ~40B active MoE
      • 1M-token context
      • Two thinking-effort levels (incl. 'Max')
      • Agentic / long-horizon coding
      • MIT license

      Innovations

      • IndexShare attention (~2.9x FLOP reduction at 1M-token context)
      • Usable 1M-token context window

      Claude Fable 5

      2026-06-09
      Anthropic
      Architecture not disclosed

      Anthropic's first publicly available Mythos-class model, priced at $10/M input and $50/M output. Launched alongside the gated Claude Mythos 5 (same base model with safeguards lifted for a small set of cyberdefenders and infrastructure providers), which is not publicly available and is not tracked here. A safety system routes a minority of sessions (~5% on average) to Claude Opus 4.8. Vendor-reported benchmarks: SWE-bench Pro 80.3, HLE 59.0 (no tools; 64.5 with tools), Terminal-Bench 2.1 88.0.

      GPQA93.18%
      HLE59%
      SWE-bench Pro80.3%
      MMLU Pro91.5%
      Agent Arena+14.2%#1

      Key Capabilities

      • Autonomous long-horizon tasks
      • Vision input
      • Agentic workflows
      • Software engineering

      Innovations

      • First publicly available Mythos-class model
      • Safety routing to Claude Opus 4.8

      Reve 2.0

      2026-06-03
      Reve
      Architecture not disclosed

      Layout-first 4K text-to-image model from Reve. Represents each element with a position, size and local description for code-like, addressable editing. Native 4096×4096 output.

      Key Capabilities

      • Layout-first generation
      • Native 4K (4096×4096) output
      • Addressable / editable elements
      • Text rendering in images

      Innovations

      • Layout-as-prompt (position + size + local description)
      • Native 4096×4096 output
      Microsoft
      Architecture not disclosed

      Microsoft's first in-house coding model, part of the MAI family unveiled at Build 2026 and built on Maia 200 silicon without distillation from other labs. Live in GitHub Copilot and VS Code, and available to developers via OpenRouter, Fireworks and Baseten with tunable weights. Vendor-reported benchmarks: SWE-bench Pro 51.2, GPQA Diamond 84.6 (model card).

      GPQA84.6%
      SWE-bench Pro51.2%

      Key Capabilities

      • Agentic coding
      • GitHub Copilot integration
      • Low token usage

      Innovations

      • First Microsoft in-house coding model
      • Trained on Maia 200 silicon
      • Developer-tunable weights

      MAI-Image-2.5

      2026-06-02
      Microsoft
      Architecture not disclosed

      Text-to-image and image-editing model from Microsoft's MAI family, unveiled at Build 2026 and built in-house on Maia 200 silicon. Available through Microsoft Foundry, with a faster Flash variant that is not tracked separately here.

      Key Capabilities

      • Text-to-image
      • Image editing

      Innovations

      • First Microsoft in-house image model
      • Trained on Maia 200 silicon

      MAI-Thinking-1

      2026-06-02
      Microsoft
      Architecture not disclosed

      Microsoft AI's first flagship reasoning model, a medium-sized model in the first-party MAI family launched at Build 2026. Distributed via OpenRouter, Fireworks and Baseten.

      GPQA84.2%
      SWE-bench Pro52.8%

      Key Capabilities

      • Extended reasoning
      • Agentic coding
      • Mathematical reasoning
      • Tool use

      Innovations

      • First-party Microsoft reasoning flagship
      • MAI model family (Build 2026)
      May

      Claude Opus 4.8

      2026-05-28
      Anthropic

      Anthropic's flagship upgrade to Opus 4.7, released at the same standard price ($5/M input, $25/M output). 1M-token input context with up to 128K output tokens. Adds a Fast mode running at 2.5x output speed ($10/M input, $50/M output) alongside the standard tier. New 'dynamic workflows' tool for Claude Code, effort control in claude.ai and Cowork, and a Messages API that now accepts system entries mid-conversation. Vendor-reported benchmarks: GPQA Diamond 93.6, SWE-bench Verified 88.6, SWE-bench Pro 69.2, Terminal-Bench 2.1 74.6. Artificial Analysis Intelligence Index v4.0: 61.

      GPQA93.6%
      SWE-bench88.6%
      SWE-bench Pro69.2%
      MMLU Pro89.58%
      Agent Arena+9.0%#2

      Key Capabilities

      • 1M-token input context window
      • Up to 128K output tokens
      • Vision input
      • Computer use
      • Agentic workflows

      Innovations

      • Dynamic Workflows (Claude Code)
      • Effort Control (claude.ai & Cowork)
      • Mid-Conversation System Entries (Messages API)
      • Fast Mode (2.5x speed tier)

      Qwen3.7-Max

      2026-05-20
      Alibaba
      Architecture not disclosed

      Proprietary flagship in Alibaba's Qwen Max line. Closed weights — the Plus variant is open-sourced while Max stays closed, consistent with Alibaba's Max/Plus split since Qwen2.5-Max. 1M-token context (up from 256K in the Qwen3.6-Max preview), text-only input and output. MoE architecture details are not officially disclosed. Previewed on Arena AI on 2026-05-14, formally launched at the 2026 Alibaba Cloud Summit on 2026-05-20. Alibaba-reported benchmarks (verified post-launch via AA Intelligence Index v4.0 reproduction): GPQA Diamond 92.4, HLE 41.4, SWE-bench Verified 80.4, Terminal-Bench 2.0 69.7. AA Intelligence Index v4.0: 57 (initial Index run reported 56.6).

      GPQA92.4%
      HLE41.4%
      SWE-bench80.4%
      MMLU Pro89.31%

      Key Capabilities

      • 1M-token context window
      • Text-only input and output
      • Closed weights

      Innovations

        Command A+

        2026-05-20
        Cohere
        218B total·25B activeMoE

        218B sparse Mixture-of-Experts (~25B active) open-weight flagship under Apache 2.0, runnable on as few as 2x H100 GPUs. Enterprise- and RAG-focused with strong multilingual coverage. Available as open weights and via Cohere's API. Early independent benchmarks (Artificial Analysis): GPQA Diamond ~76, HLE ~11, MMMU-Pro 63, Intelligence Index ~37 — flagged as estimates, AA's full evaluation is still ongoing.

        GPQA76%
        HLE11%

        Key Capabilities

        • 218B sparse MoE / ~25B active
        • Apache 2.0 open weights
        • Multilingual
        • Enterprise / RAG focus
        • Runs on 2x H100

        Innovations

        • Apache 2.0 open-weight enterprise flagship
        Google
        Architecture not disclosed

        Fast-tier model in the Gemini 3.5 generation, announced at Google I/O 2026. GA at launch and set as the default model in the Gemini app and AI Mode in Google Search. Architecture is not publicly disclosed. Vendor-reported benchmarks: GPQA Diamond 90.4, MMMU-Pro 81.2, SWE-bench Verified 78. Independent (Artificial Analysis, ~5 days post-launch): HLE 40.2 (exact match to vendor figure), Terminal-Bench 2.1 76.2, AA Intelligence Index v4.0: 55. Distinct from 'Gemini Omni', a separate video/world model announced at the same event.

        GPQA90.4%
        HLE40.2%
        SWE-bench78%
        MMLU Pro89.52%
        Agent Arena+0.0%#15

        Key Capabilities

        • Default model in Gemini app and Search AI Mode
        • GA at launch (2026-05-19)
        • Multimodal input

        Innovations

        • Output speed 207.9 tok/s — AA Speed rank #2/148 at launch (Speed is the product identity behind the Flash name)

        ERNIE 5.1

        2026-05-08
        Baidu
        800B total·36B activeMoE

        MoE flagship derived from ERNIE 5.0 via 'multi-dimensional elastic pre-training' as an optimal sub-network. Parameter counts are estimates (~1/3 of 5.0's total, ~1/2 of 5.0's active params); Baidu has not published exact figures. Cloud-only via Baidu Qianfan, no Western public API. Multimodal (text/image/audio/video) inherited from 5.0. Baidu claims ~6% of comparable training cost — unverified by third parties. No standalone technical report published. Self-reported benchmarks; non-EU hosting.

        Key Capabilities

        • Multimodal (text/image/audio/video)
        • 128k context window
        • MoE sub-network of ERNIE 5.0
        • Cloud-only (Baidu Qianfan)
        • Closed weights

        Innovations

        • Multi-Dimensional Elastic Pre-Training (sub-network extraction from ERNIE 5.0)
        • ~6% of comparable training cost (Baidu claim, unverified)

        GPT-5.5 Instant

        2026-05-05
        OpenAI
        Architecture not disclosed

        Default ChatGPT model, replacing GPT-5.3 Instant. API identifier 'chat-latest'. Hallucination-focused update with material gains on math and multimodal reasoning vs. its predecessor. Architecture not publicly disclosed.

        Key Capabilities

        • Default ChatGPT model
        • API identifier: chat-latest
        • Multimodal (text + image)
        • Reduced hallucinations on high-stakes prompts
        • Tool use / function calling

        Innovations

        • ChatGPT default for hundreds of millions of users
        • 52.5% Hallucination Reduction vs 5.3 Instant (high-stakes, internal eval)
        • AIME 2025 Jump 65.4 to 81.2 vs 5.3 Instant
        • Low-latency 'Instant' tier — OpenAI states GPT-5.5 matches GPT-5.4 per-token latency in real-world serving (no AA Instant entry yet)
        Apr
        Mistral
        128B

        128B dense flagship with 256k context, multimodal (text + vision), open weights under a modified MIT license. First Mistral flagship trained on the new 13,800-GPU Paris facility. Merges the Magistral reasoning line and the Devstral 2 coding line into one model. Mistral did not publish GPQA/HLE/MMLU-Pro at launch; granular third-party scores still pending. Independent composite available: AA Intelligence Index v4.0 = 39 (#2 in class), with very verbose output (~90M tokens vs. 16M class average).

        SWE-bench77.6%

        Key Capabilities

        • 128B dense parameters
        • 256k context window
        • Multimodal (text + vision)
        • Reasoning + coding unified
        • Open weights (Modified MIT)
        • Self-hostable on 4 GPUs (~70GB VRAM Q4)

        Innovations

        • First flagship trained on Mistral's 13,800-GPU Paris facility
        • Unified Magistral (reasoning) + Devstral 2 (coding) in one model
        • Modified MIT release at flagship scale

        Grok 4.3

        2026-04-30
        X.AI
        Architecture not disclosed

        Reasoning-first flagship with a 1M-token context window and native video input. Agentic-focused release at lower pricing than Grok 4.1.

        GPQA90.1%
        MMLU Pro85.84%
        Agent Arena-7.2%#22

        Key Capabilities

        • 1M context window
        • Native video input
        • Reasoning-first

        Innovations

          MiMo-V2.5-Pro

          2026-04-27
          Xiaomi
          1.0T total·42B activeMoE

          Xiaomi's flagship coding/agentic open-weights model. 1.02T total / 42B active hybrid MoE (70 layers, 384 routed experts, 8 active per token) with interleaved Sliding Window + Global Attention at a 6:1 ratio and 128-token window for long-context efficiency. 1M context, MIT licensed, day-0 SGLang/vLLM support. Reports SWE-bench Pro 57.2, GDPVal-AA Elo 1581, ClawEval 64% Pass³ at ~70K tokens per trajectory, and a SysY compiler-in-Rust task solved 233/233 in 4.3h with 672 tool calls. Project lead Fuli Luo (ex-DeepSeek).

          SWE-bench Pro57.2%
          MMLU Pro84.59%

          Key Capabilities

          • 1.02T total / 42B active (hybrid MoE)
          • 1M context window
          • Hybrid SWA + Global Attention (6:1 ratio, 128-token window)
          • FP8 E4M3 mixed precision
          • Day-0 SGLang/vLLM support
          • MIT license

          Innovations

          • Interleaved Sliding Window + Global Attention (6:1)
          • Lightweight MTP modules with dense FFNs
          • Long-horizon agent trajectories (~70K tokens / 672 tool calls)

          DeepSeek-V4-Pro

          2026-04-24
          Deep Seek
          1.6T total·49B activeMoE

          Preview release, available via API and open weights. 1.6T total / 49B active MoE with hybrid attention and manifold-constrained hyper-connections. Reports 27% of single-token inference FLOPs and 10% of KV cache vs DeepSeek-V3.2 in a 1M-token setting.

          GPQA90.1%
          HLE37.7%
          SWE-bench80.6%
          SWE-bench Pro55.4%
          MMLU Pro87.5%

          Key Capabilities

          • 1.6T total / 49B active MoE
          • 1M token context window
          • Three reasoning effort modes (incl. Think Max)
          • MIT license

          Innovations

          • Hybrid attention: CSA + HCA
          • Manifold-Constrained Hyper-Connections
          • 27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens
          Deep Seek
          284B total·13B activeMoE

          Preview release, available via API and open weights; superseded by DeepSeek-V4-Flash-0731 on 2026-07-31. Smaller 284B total / 13B active MoE sibling of V4-Pro, sharing hybrid attention and hyper-connection architecture.

          GPQA88.1%
          HLE34.8%
          SWE-bench79%
          SWE-bench Pro52.6%
          MMLU Pro86.2%

          Key Capabilities

          • 284B total / 13B active MoE
          • 1M token context window
          • Three reasoning effort modes (incl. Think Max)
          • MIT license

          Innovations

          • Hybrid attention: CSA + HCA
          • Manifold-Constrained Hyper-Connections
          • Shared architecture with V4-Pro at smaller scale

          GPT-5.5

          2026-04-23
          OpenAI
          Architecture not disclosed

          Agentic-focused flagship with 1M context window (400K in Codex). Thinking and Pro variants available. Codename 'Spud'.

          SWE-bench Pro58.6%
          MMLU Pro88.14%
          Agent Arena+8.3%#3

          Key Capabilities

          • 1M context window
          • Thinking mode
          • Pro variant ($30/$180 per M tokens)
          • 400K context in Codex
          • Agentic workflows
          • API pricing: $5/$30 per M tokens (base)

          Innovations

          • 1M Context Window (largest OpenAI flagship)
          • Codex-Optimized 400K Context
          • Thinking + Pro Variant Lineup
          OpenAI
          Architecture not disclosed

          First OpenAI image model with reasoning capabilities. 2K resolution, multilingual text rendering, up to 8 images per prompt, web search integration. API name 'gpt-image-2'. Replaces DALL-E 3 (deprecated May 12, 2026).

          Key Capabilities

          • Text-to-Image
          • 2K resolution output
          • Reasoning-guided generation
          • Multilingual text rendering (Japanese, Korean, Hindi, Bengali)
          • Up to 8 images per prompt
          • Web search integration
          • API: gpt-image-2 ($8/$30 per M tokens)

          Innovations

          • Reasoning-Guided Image Generation
          • Multilingual Text Rendering
          • Web Search Integration for Image Context
          • DALL-E 3 Successor

          GR00T N1.7

          2026-04-17
          NVIDIA
          3B

          Open vision-language-action model for humanoid robots, released in early access under Apache 2.0 — the first fully commercially licensed model of the GR00T line. A vision-language backbone is paired with a diffusion transformer head that denoises continuous actions via flow matching. NVIDIA credits the generalization and language-following gains over N1.6 to 20,000 hours of EgoScale human egocentric video in pretraining. At 3B parameters it is small enough to run on local hardware, which is why it also appears in the Local tab.

          Key Capabilities

          • Vision-language-action robot control
          • Flow-matching action transformer head
          • Humanoid whole-task policies
          • Apache 2.0 open weights

          Innovations

          • Open, commercially licensed humanoid VLA
          • 20K hours of EgoScale human video in pretraining

          Claude Opus 4.7

          2026-04-16
          Anthropic

          Anthropic's most capable generally available model. Default in Claude Code. Dense architecture with new tokenizer, file-system-based cross-session memory, new 'xhigh' effort level, and task budgets (public beta) for agentic loops. 1M context window with no long-context premium. Sits below Mythos Preview in capability, above Opus 4.6. $5/M input, $25/M output tokens.

          GPQA94.2%
          SWE-bench87.6%
          SWE-bench Pro64.3%
          MMLU Pro89.87%
          Agent Arena+8.1%#4

          Key Capabilities

          • 1M context window
          • Vision input up to 3.75 MP
          • Computer use with 1:1 pixel mapping
          • Agentic workflows
          • Cross-session file-system memory

          Innovations

          • xhigh Effort Level
          • Task Budgets (Public Beta)
          • File-System Memory
          • New Tokenizer

          Qwen3.6-35B-A3B

          2026-04-16
          Alibaba
          35B total·3B activeMoE

          Successor family to Qwen3.5 from the Qwen team, focused on agentic coding. 35B total / 3B active MoE with 262K context window and thinking preservation across conversation turns. Compatible with OpenClaw, Claude Code, and Cline. Apache 2.0 weights on HuggingFace.

          Key Capabilities

          • 35B total / 3B active MoE
          • 262K context window
          • Agentic coding focus
          • Repository-level reasoning
          • Thinking preservation across turns
          • Apache 2.0

          Innovations

          • Thinking Preservation Across Conversation
          • Repository-Level Reasoning
          • OpenClaw/Claude Code/Cline Compatible

          π0.7

          2026-04-16
          Physical Intelligence
          Architecture not disclosed

          Steerable robot foundation model from Physical Intelligence, presented as the point where VLAs start generalising rather than replaying training data. Physical Intelligence reports 82.1% success on trained tasks and 47.3% zero-shot on unseen ones across seven robot embodiments and 50+ manipulation tasks, and demonstrates the model operating an air fryer it had only encountered in two fragmentary training episodes. Weights are not public.

          Key Capabilities

          • Vision-language-action robot control
          • Zero-shot spatial generalization
          • Semantic steering from language alone
          • Multi-step compositional manipulation

          Innovations

          • Emergent task generalization beyond the training distribution
          • One policy across seven robot embodiments

          Muse Spark

          2026-04-08
          Meta
          Architecture not disclosed

          First model from Meta Superintelligence Labs under Alexandr Wang. Powers Meta AI across Facebook, Instagram, WhatsApp, Messenger and Ray-Ban glasses. Natively multimodal (text/image/voice input, text output) with multi-agent 'Contemplating mode'. Closed-source departure from the Llama lineage. Free on meta.ai and the Meta AI app; API in private preview. Consumer-focused, strong in health and multimodal; coding gap acknowledged by Meta.

          GPQA89.65%
          HLE39.9%
          MMLU Pro87.32%

          Key Capabilities

          • Multimodal perception
          • Voice input
          • Reasoning and agentic tasks
          • Powers Meta AI products
          • Ray-Ban smart glasses integration
          • Health Q&A

          Innovations

          • Meta Superintelligence Labs Debut
          • Contemplating Mode (multi-agent)
          • Closed-source Departure from Llama
          Anthropic
          Architecture not disclosed

          Anthropic's 'Capybara tier' flagship, withheld from public release due to cybersecurity risk. Available only to 12 Project Glasswing enterprise partners (incl. Amazon, Apple, Google, Microsoft, NVIDIA) via a $100M credit pool. No public API. $25/M input, $125/M output tokens.

          GPQA94%
          HLE56.8%
          SWE-bench93.9%
          SWE-bench Pro77.8%

          Key Capabilities

          • Frontier reasoning
          • Advanced coding
          • Agentic workflows
          • Mathematical reasoning

          Innovations

          • Capybara Tier
          • Project Glasswing Restricted Access
          • Withheld for Cybersecurity Risk

          Happy Horse 1.0

          2026-04-07
          Alibaba
          15B

          Video generation model from Alibaba's ATH AI Innovation Unit. 15B-parameter single-stream Transformer (40 layers) with native audio/video/text token fusion. #1 in both text-to-video and image-to-video on Artificial Analysis Video Arena (Elo blind test). Initially released anonymously; confirmed as Alibaba on April 10, 2026. Beta — no public API or released weights.

          Key Capabilities

          • Text-to-video
          • Image-to-video
          • Native audio/video/text token fusion

          Innovations

          • Single-stream 40-layer Transformer
          • Video Arena #1 (T2V and I2V)
          • Native Multimodal Token Fusion
          Mar

          Qwen3.6 Plus

          2026-03-31
          Alibaba
          Architecture not disclosed

          Flagship cloud API model with a 1M-token context window and always-on chain-of-thought reasoning. Agentic coding focus.

          GPQA88.2%
          HLE25.7%
          SWE-bench78.8%
          MMLU Pro87.67%
          Agent Arena-4.2%#20

          Key Capabilities

          • 1M context window
          • Always-on chain-of-thought
          • Agentic coding

          Innovations

            Midjourney
            Architecture not disclosed

            Major architectural overhaul on rewritten codebase. ~5x faster generation than V7, native 2K resolution, significantly improved text rendering, better prompt adherence for complex multi-element compositions.

            Key Capabilities

            • ~5x faster generation
            • Native 2K resolution (--hd)
            • Improved text rendering
            • Better prompt adherence
            • New --q 4 quality mode

            Innovations

            • Rewritten Codebase
            • Native 2K Resolution
            • Enhanced Text Rendering
            • Scene Coherence Mode

            GPT-5.4 mini

            2026-03-17
            OpenAI
            Architecture not disclosed

            Fast reasoning model approaching GPT-5.4 performance at 2x the speed, optimized for sub-agent workloads.

            GPQA88%
            HLE28.2%
            SWE-bench Pro54.4%
            MMLU Pro84.55%

            Key Capabilities

            • Reasoning (configurable effort)
            • Multimodal (text + image)
            • Tool use / function calling
            • Computer use
            • 2x faster than GPT-5 mini

            Innovations

            • Sub-Agent Architecture Optimized
            • Near-Flagship Performance at Mini Scale
            • 2x Speed vs GPT-5 mini

            GPT-5.4 nano

            2026-03-17
            OpenAI
            Architecture not disclosed

            Smallest, cheapest GPT-5.4 variant for classification, extraction, and cost-sensitive sub-agent tasks.

            GPQA82.8%
            HLE24.3%
            SWE-bench Pro52.4%

            Key Capabilities

            • Multimodal (text + image)
            • Tool use / function calling
            • Classification & extraction
            • Cost-optimized ($0.20/1M input)

            Innovations

            • Ultra Cost-Efficient Reasoning
            • Sub-Agent Era Design
            • Smallest GPT-5.4 Variant

            Mistral Small 4

            2026-03-16
            Mistral
            119B total·6B activeMoE

            119B MoE model (6B active) unifying reasoning, vision, and coding with configurable reasoning effort.

            GPQA71.2%
            HLE9.5%
            MMLU Pro78%

            Key Capabilities

            • 119B MoE (6B active)
            • Configurable reasoning effort
            • Multimodal (vision)
            • Agentic coding
            • 256k context
            • Apache 2.0

            Innovations

            • Unified Magistral + Pixtral + Devstral
            • 128-Expert MoE Architecture
            • 40% Latency Reduction vs Small 3
            • 3x Throughput vs Small 3

            GPT-5.3 Instant

            2026-03-05
            OpenAI
            Architecture not disclosed

            Fast everyday model with reduced hallucinations and improved web search.

            Key Capabilities

            • Reduced hallucinations
            • Improved web search
            • Natural conversation
            • Creative writing

            Innovations

            • 26.8% Hallucination Reduction (web)
            • 19.7% Hallucination Reduction (internal)
            • Balanced Reasoning + Search
            OpenAI
            Architecture not disclosed

            Multi-step reasoning model with native computer-use and 1M token context.

            GPQA92.8%
            HLE42%
            SWE-bench Pro57.7%
            MMLU Pro87.48%
            Agent Arena+6.5%#9

            Key Capabilities

            • Multi-step reasoning
            • Native computer use
            • 1M token context (API)
            • Tool-heavy workflows
            • Agentic tasks

            Innovations

            • Native Computer Use
            • Tool Search
            • Efficient Token Usage
            • 28-point OSWorld Jump

            GPT-5.4 Pro

            2026-03-05
            OpenAI
            Architecture not disclosed

            Highest-capability model optimized for quality and depth over speed.

            GPQA94.4%

            Key Capabilities

            • Highest capability
            • Decision-ready outputs
            • Extended reasoning
            • Native computer use
            • 1M token context (API)

            Innovations

            • ARC-AGI-2 Leader (83.3%)
            • Quality over Speed Optimization
            Feb

            Nano Banana 2

            2026-02-26
            Google
            Architecture not disclosed

            Pro-quality AI image generation at Flash speed, combining Nano Banana Pro features with Gemini Flash performance.

            Key Capabilities

            • 4K image generation
            • Text rendering in images
            • Real-time knowledge integration
            • Multi-character consistency
            • Extreme aspect ratios

            Innovations

            • Gemini 3.1 Flash Image
            • Web Search Integration
            • Multi-language Text Rendering
            • Subject Consistency (up to 5 characters)
            Google
            Architecture not disclosed

            Google DeepMind's Nano Banana 2 (API gemini-3.1-flash-image-preview), third in the Nano Banana family after the original (Gemini 2.5 Flash, Aug 2025) and Nano Banana Pro (Gemini 3 Pro, Nov 2025). Brings Pro-tier image quality to the Flash architecture at roughly half the per-image API price (~$0.067 vs $0.134 for a 2K image).

            Key Capabilities

            • Text-to-Image
            • Pro-quality at Flash speed
            • Strong text rendering
            • Lower per-image cost

            Innovations

            • Pro Quality at Flash Speed
            • Half-Price 2K Images

            Gemini 3.1 Pro

            2026-02-19
            Google
            Architecture not disclosed

            Major reasoning upgrade with record GPQA and 2x ARC-AGI-2 improvement over Gemini 3 Pro.

            GPQA94.3%
            HLE44.4%
            SWE-bench80.6%
            SWE-bench Pro54.2%
            MMLU92.6%
            MMLU Pro90.99%
            Agent Arena-0.8%#17

            Key Capabilities

            • 1M context window
            • 64K output
            • Natively multimodal
            • Agentic workflows

            Innovations

            • ARC-AGI-2 Regime Change
            • Deep Think Reasoning
            • SVG Animation Generation
            Anthropic
            Architecture not disclosed

            Near-flagship performance at mid-tier pricing with 1M context window and advanced computer use.

            GPQA89.9%
            HLE49%
            SWE-bench79.6%
            MMLU Pro87.34%
            Agent Arena+3.2%#12

            Key Capabilities

            • 1M context window
            • Advanced computer use
            • Adaptive thinking
            • Financial analysis

            Innovations

            • Flagship-tier at Mid-tier Pricing
            • ARC-AGI-2 Leap
            • OSWorld Computer Use

            Qwen3.5

            2026-02-16
            Alibaba
            397B total·17B activeMoE

            Major architectural upgrade over Qwen3. 397B MoE (17B active) flagship with 1M context, natively multimodal, 8-19x higher decoding throughput. Covers 200+ languages.

            GPQA72%
            HLE18%
            SWE-bench48.5%
            MMLU90.2%
            MMLU Pro87.18%

            Key Capabilities

            • 397B total / 17B active MoE (flagship)
            • 1M context window
            • Natively multimodal
            • Reasoning by default
            • 200+ languages
            • Apache 2.0

            Innovations

            • 8-19x Throughput vs Qwen3-Max
            • Native Multimodal All Sizes
            • 122B-A10B Runs on MacBook 64GB
            ByteDance

            Frontier-Flaggschiff der Seed-2.0-Familie (Pro/Lite/Mini/Code) von ByteDance, 256k Kontext, Schwerpunkt Reasoning, Coding und Agent-Tasks. Treibt die Doubao-App, Chinas meistgenutzten KI-Chatbot. Proprietär über Volcano Engine.

            GPQA88.9%
            SWE-bench76.5%
            MMLU Pro87%

            Key Capabilities

            • 256k Context
            • Reasoning / Agent
            • Coding
            • Multilingual

            Innovations

            • Seed-2.0-Foundation-Familie
            • Volcano-Engine-API

            Seedance 2.0

            2026-02-12
            ByteDance
            Architecture not disclosed

            ByteDance text/image/audio-to-video model with unified audio-video joint generation. Generates 4-15s clips up to 1080p across multiple aspect ratios. Available via Volcano Engine and API.

            Key Capabilities

            • Text-to-video
            • Image-to-video
            • Audio-video joint generation
            • Up to 1080p output
            • 4-15s clips

            Innovations

            • Unified multimodal audio-video architecture
            • Multimodal reference inputs (text/image/audio/video)
            • Director-level camera & lighting control

            GLM-5

            2026-02-11
            Zhipu AI
            744B total·40B activeMoE

            744B MoE (40B active) open-weight model with DeepSeek Sparse Attention and strong agentic/frontend coding performance.

            MMLU Pro86.03%

            Key Capabilities

            • 744B total / 40B active MoE
            • 200K context
            • Agentic workflows
            • Frontend coding (98% build success)
            • MIT license

            Innovations

            • DeepSeek Sparse Attention
            • Slime Async RL Framework
            • 98% Frontend Build Success Rate
            • Stealth-Launched as Pony Alpha on OpenRouter

            GPT-5.3 Codex

            2026-02-05
            OpenAI
            Architecture not disclosed

            Agentic coding model, 25% faster, first model instrumental in creating itself.

            SWE-bench Pro56.8%

            Key Capabilities

            • Agentic coding
            • 25% faster
            • Self-developing
            • Tool use

            Innovations

            • Self-Development
            • Terminal-Bench SOTA
            • OSWorld Leader

            Claude Opus 4.6

            2026-02-05
            Anthropic
            Architecture not disclosed

            Next-generation flagship model with adaptive thinking and agent teams.

            GPQA89.2%
            HLE47.3%
            SWE-bench Pro56.8%
            MMLU95.2%
            MMLU Pro89.11%
            Agent Arena+6.7%#8

            Key Capabilities

            • Adaptive thinking
            • 1M context window
            • Agent teams
            • 128K output

            Innovations

            • Agent Teams
            • Adaptive Reasoning
            • Claude Cowork

            Kling 3.0

            2026-02-05
            Kuaishou
            Architecture not disclosed

            Latest flagship with AI Director capability and multi-shot storyboard.

            Key Capabilities

            • 15s video
            • Native audio
            • 2K/4K images
            • AI Director
            • Video generation

            Innovations

            • AI Director
            • Multi-shot Storyboard
            • Character/Voice Replication
            Jan
            X.AI
            Architecture not disclosed

            High-definition text-to-video model with synchronized cinematic audio.

            Key Capabilities

            • Text-to-video
            • Native audio
            • Video editing
            • HD output

            Innovations

            • End-to-end video+audio generation
            • #1 Artificial Analysis video ranking
            • Cinematic quality
            Alibaba
            Architecture not disclosed

            Trillion-parameter reasoning model with adaptive tool use and test-time scaling.

            GPQA87.4%
            HLE36.5%
            SWE-bench75.3%
            MMLU Pro84.98%

            Key Capabilities

            • Advanced reasoning
            • Adaptive tool use
            • Test-time scaling
            • Code interpretation
            • Web search integration
            • Video understanding

            Innovations

            • Reasoning Mode
            • Adaptive Tool Invocation
            • 1T+ MoE Architecture
            • Dynamic Compute Allocation

            Kimi K2.5

            2026-01-27
            Moonshot AI
            1T total·32B activeMoE

            1T MoE (32B active) open-weight model with native multimodal vision and agent swarm support for up to 100 sub-agents.

            MMLU Pro85.91%

            Key Capabilities

            • 1T total / 32B active MoE
            • 262K context
            • Natively multimodal (MoonViT 400M)
            • Agent swarms (100 sub-agents)
            • 1,500 parallel tool calls
            • Modified MIT license

            Innovations

            • MoonViT Vision Encoder (400M params)
            • 100 Sub-Agent Swarm Support
            • 1,500 Parallel Tool Calls
            • Top-Ranked Artificial Analysis & LMArena at Launch

            ERNIE 5.0

            2026-01-22
            Baidu
            2.4T total·72B activeMoE

            2.4T-parameter ultra-sparse MoE flagship in public preview via ERNIE Bot. Natively omnimodal (text/image/audio/video jointly modeled from pretraining). #1 Chinese model on LMArena Text and Vision (#8 globally). China-market only; no Western public API.

            MMLU85%

            Key Capabilities

            • Text
            • Image
            • Audio
            • Video
            • Omnimodal reasoning
            • Agentic workflows

            Innovations

            • Natively Omnimodal Unified Autoregressive Framework
            • Ultra-Sparse MoE (<3% activation per inference)
            • PaddlePaddle Framework
            2025
            Dec

            Gemini 3 Flash

            2025-12-17
            Google
            Architecture not disclosed

            Flash-tier model offering frontier-level reasoning at lower cost and higher speed. Predecessor to Gemini 3.5 Flash.

            GPQA90.4%
            HLE33.7%
            SWE-bench78%
            MMLU Pro88.59%
            Agent Arena-8.5%#24

            Key Capabilities

            • High-speed reasoning
            • Multimodal input

            Innovations

              GPT Image 1.5

              2025-12-16
              OpenAI
              Architecture not disclosed

              OpenAI's second-generation flagship image model (API gpt-image-1.5), rolled out globally to all ChatGPT tiers. About 4x faster generation, stronger instruction adherence, edit precision that preserves lighting and composition, and improved text rendering for dense typography. The high-fidelity variant is the strongest of the GPT Image 1.x line on the LMArena Text-to-Image arena.

              Key Capabilities

              • Text-to-Image
              • High-fidelity generation
              • ~4x faster generation
              • Instruction-preserving edits
              • Dense text rendering

              Innovations

              • High-Fidelity Editing
              • 4x Faster Generation

              GWM-1

              2025-12-11
              Runway
              Architecture not disclosed

              Runway's first General World Model: an autoregressive model built on top of Gen-4.5 that generates frame by frame, runs in real time at 24 fps and is steered interactively through camera pose, robot commands or audio. Ships as three variants, GWM Worlds for explorable environments, GWM Avatars for speaking characters with facial expression and lip sync, and GWM Robotics for synthetic robot training data, which Runway plans to merge into one model. Announced together with native audio for Gen-4.5 and pitched as more general than Google's Genie 3: a simulator for training agents in domains such as robotics and life sciences.

              Key Capabilities

              • Real-time frame-by-frame world simulation at 24 fps
              • Interactive control via camera pose, robot commands and audio
              • Explorable worlds, avatars and robotics variants
              • Synthetic training data for robots

              Innovations

              • Autoregressive world model built on a video generation model (Gen-4.5)
              • Action-conditioned real-time simulation

              Seedream 4.5

              2025-12-04
              ByteDance
              Architecture not disclosed

              All-round upgrade to Seedream 4.0 with stronger in-image text rendering, roughly 10× faster generation than 4.0, and multi-reference fusion accepting up to 10 reference images. Output up to 2048×2048, offered as text-to-image and image-edit variants.

              Key Capabilities

              • Text-to-Image
              • Image editing
              • Strong in-image text rendering
              • Multi-reference fusion (up to 10 images)
              • Up to 2048×2048 output

              Innovations

              • ~10× faster than Seedream 4.0
              • Multi-reference fusion
              • Improved text rendering

              Kling 2.6

              2025-12-03
              Kuaishou
              Architecture not disclosed

              First simultaneous audio-visual generation model.

              Key Capabilities

              • Audio-visual generation
              • Speech & dialogue
              • 1080p/10s
              • Video generation

              Innovations

              • Simultaneous Audio-Visual
              • Multi-language Audio
              • Ambient Sound Generation
              Mistral
              675B totalMoE

              Flagship Mistral Large 3 (675B MoE) and efficient Ministral 3 edge models.

              Key Capabilities

              • 675B MoE
              • Multimodal
              • 256k Context
              • Edge-optimized Ministral

              Innovations

              • Native Multimodal MoE
              • Edge-Cloud Model Family
              • Apache 2.0 Weights

              DeepSeek-V3.2

              2025-12-01
              Deep Seek
              671B total·37B activeMoE

              Successor to DeepSeek V3. 671B MoE (37B active) with DeepSeek Sparse Attention for long-context efficiency and large-scale agentic tool-use pipeline.

              MMLU Pro84.92%

              Key Capabilities

              • 671B total / 37B active MoE
              • 164K context
              • DeepSeek Sparse Attention (DSA)
              • Agentic tool-use (1,800+ environments)
              • Strong reasoning without thinking mode
              • MIT license

              Innovations

              • DeepSeek Sparse Attention (DSA)
              • Large-Scale Task Synthesis Pipeline
              • 1,800+ Agentic Environments

              Gen-4.5

              2025-12-01
              Runway
              Architecture not disclosed

              Advanced AI video generation model with physics-based realism and director controls.

              Key Capabilities

              • 1080p Cinematic Video
              • Physics Simulation
              • Director Mode 2.0
              • Motion Brush 3.0

              Innovations

              • General World Model
              • Physics-based Realism
              • Latent Diffusion Transformers
              Nov

              Flux 2.0

              2025-11-25
              Black Forest Labs
              Architecture not disclosed

              Next generation with improved photorealism and typography.

              Key Capabilities

              • Image reference
              • Photorealism
              • Typography
              • Prompt understanding

              Innovations

              • Flux.2 Pro
              • Flux.2 Dev
              • Apache 2.0 License
              • 32B Open Source

              Claude Opus 4.5

              2025-11-24
              Anthropic
              Architecture not disclosed

              Next-generation flagship model with breakthrough reasoning and research capabilities.

              GPQA93.2%
              HLE44.8%
              SWE-bench76.9%
              MMLU94.5%
              MMLU Pro87.26%

              Key Capabilities

              • Advanced reasoning
              • Scientific research
              • Long-context understanding
              • Multi-step planning

              Innovations

              • Enhanced Constitutional AI
              • Extended Context Processing
              • Advanced Tool Integration
              Google
              Architecture not disclosed

              Advanced image generation model building upon Nano Banana.

              Key Capabilities

              • Text-to-Image
              • High fidelity generation
              • Multimodal integration

              Innovations

              • Nano Banana Enhancement
              • Advanced Image Synthesis
              • Gemini 3 Integration
              OpenAI
              Architecture not disclosed

              Specialized ultra-high-performance coding model with breakthrough architecture.

              GPQA89.7%
              HLE38.2%
              SWE-bench75.4%
              MMLU93.8%

              Key Capabilities

              • Advanced code generation
              • Multi-language mastery
              • System-level reasoning

              Innovations

              • Neural Code Synthesis
              • Real-time Code Optimization
              • Integrated Debugging

              Gemini 3

              2025-11-18
              Google
              Architecture not disclosed

              Next major generation with agentic-first architecture.

              GPQA91.9%
              HLE41.2%
              SWE-bench72.5%
              MMLU93.1%
              MMLU Pro90.1%

              Key Capabilities

              • Agentic-first
              • Infinite context
              • Personalization

              Innovations

              • Antigravity IDE
              • Native Agentic Workflow

              Grok 4.1

              2025-11-17
              X.AI
              Architecture not disclosed

              Refined Grok 4 with improved safety and alignment.

              GPQA89.3%
              HLE36.7%
              SWE-bench62.1%
              MMLU91.2%

              Key Capabilities

              • Truth seeking
              • Reduced bias
              • Enhanced safety

              Innovations

                ChatGPT 5.1

                2025-11-12
                OpenAI
                Architecture not disclosed

                Major update with "Instant" and "Thinking" modes for adaptive reasoning.

                GPQA86.6%
                HLE35.4%
                SWE-bench67.8%
                MMLU92.3%

                Key Capabilities

                • Adaptive reasoning
                • Instant/Thinking modes
                • Personalization

                Innovations

                • Node-based Agent Builder

                Marble

                2025-11-12
                World Labs
                Architecture not disclosed

                World Labs' first commercial product and, in its own words, a frontier multimodal world model. Text prompts, photos, videos, panoramas or coarse 3D layouts become persistent, editable 3D environments that export as Gaussian splats, meshes or video. The distinction from frame-by-frame world generators is that the scene geometry is fixed once generated: a world can be revisited, edited and walked through again instead of being re-dreamed on every move. Available to everyone from launch in free and paid tiers; Marble 1.1 later added auto-expanding worlds, and the World API followed in January 2026.

                Key Capabilities

                • Text, image, video and panorama to 3D world
                • Persistent, editable 3D environments
                • Export as Gaussian splats, meshes or video
                • Free and paid tiers

                Innovations

                • Persistent world geometry instead of per-frame regeneration
                • First commercial world model product
                Oct

                Odyssey-2

                2025-10-27
                Odyssey
                Architecture not disclosed

                Odyssey's general-purpose world model, trained on a large corpus of general video and interaction data rather than on game engines or synthetic environments, which the company presents as evidence that basic physics, dynamics and behaviours can be learned from video alone. It generates interactive video that responds to typed input within about fifty milliseconds and keeps running for minutes rather than seconds; Odyssey describes the interaction like using a language model: you type, and the video responds. Publicly usable at experience.odyssey.ml, later extended by Odyssey-2 Pro (API access for embedding continuous simulations into applications) and Odyssey-2 Max.

                Key Capabilities

                • Interactive video generation from typed input
                • Response to interaction in about fifty milliseconds
                • Multi-minute simulations
                • API access via Odyssey-2 Pro

                Innovations

                • World model trained on general video and interaction data rather than synthetic environments

                Veo 3.1

                2025-10-15
                Google
                Architecture not disclosed

                Refined version with enhanced realism and prompt adherence.

                Key Capabilities

                • Enhanced realism
                • Prompt adherence
                • Rich audio
                • Creative controls

                Innovations

                • Veo 3.1 Fast
                • Advanced Controls
                • Improved Audio
                Anthropic
                Architecture not disclosed

                Ultra-fast, edge-capable model.

                GPQA75.2%
                HLE21.3%
                SWE-bench51.4%
                MMLU87.2%

                Key Capabilities

                • On-device potential
                • Sub-10ms latency
                • Cost efficiency

                Innovations

                  Sep
                  Google
                  Architecture not disclosed

                  Vision-language-action model from Google DeepMind that turns camera input and natural-language instructions into robot motor commands. Shipped as a pair: Gemini Robotics 1.5 executes, while Gemini Robotics-ER 1.5 does the embodied reasoning — it can call digital tools such as web search to plan a task before handing execution over. ER 1.5 is available to developers through the Gemini API in Google AI Studio; the VLA itself went to selected partners only.

                  Key Capabilities

                  • Vision-language-action robot control
                  • Embodied reasoning with tool use (ER 1.5)
                  • Motion transfer across robot embodiments
                  • Thinks before acting

                  Innovations

                  • Split VLA / embodied-reasoning architecture
                  • Web search as a planning step for physical tasks

                  Kling 2.5 Turbo

                  2025-09-23
                  Kuaishou
                  Architecture not disclosed

                  Faster, cheaper generation with improved motion and style consistency.

                  Key Capabilities

                  • 40% faster
                  • High-motion sequences
                  • Style consistency
                  • Video generation

                  Innovations

                  • Turbo Mode
                  • 30% Cost Reduction
                  • Real-world Physics
                  Deep Seek
                  Architecture not disclosed

                  Experimental model pushing the boundaries of open weights.

                  GPQA83.7%
                  HLE29.5%
                  SWE-bench58.3%
                  MMLU89.1%

                  Key Capabilities

                  • Experimental arch
                  • Community research
                  • Novel attention

                  Innovations

                    GPT-5-codex

                    2025-09-15
                    OpenAI
                    Architecture not disclosed

                    Specialized coding model with architectural breakthroughs.

                    GPQA85.6%
                    HLE32.1%
                    SWE-bench64.2%
                    MMLU91.8%

                    Key Capabilities

                    • Architectural breakthrough
                    • Self-improving code
                    • System design

                    Innovations

                    • Integration into VS-Code

                    Sora 2

                    2025-09-15
                    OpenAI
                    Architecture not disclosed

                    Second generation video model with synchronized audio and enhanced realism.

                    Key Capabilities

                    • Synchronized audio
                    • Enhanced realism
                    • Multi-shot control
                    • Up to 20s videos

                    Innovations

                    • Native Audio Synthesis
                    • Physical Realism
                    • World-state Persistence

                    Seedream 4.0

                    2025-09-09
                    ByteDance
                    Architecture not disclosed

                    Next-generation multimodal image model that handles generation and instruction-based editing in one system, with high-definition output up to 4K.

                    Key Capabilities

                    • Text-to-Image
                    • Instruction-based image editing
                    • Multimodal image tasks
                    • Up to 4K output

                    Innovations

                    • Unified generation + editing
                    • Up to 4K resolution
                    Anthropic
                    Architecture not disclosed

                    Balanced model with next-gen capabilities.

                    GPQA81.6%
                    HLE30.5%
                    SWE-bench59.1%
                    MMLU89.8%
                    MMLU Pro87.36%

                    Key Capabilities

                    • Real-time collaboration
                    • Voice mode
                    • Advanced tool use

                    Innovations

                      Aug

                      Grok Code Fast

                      2025-08-28
                      X.AI
                      Architecture not disclosed

                      Specialized high-speed coding model.

                      GPQA79.5%
                      HLE22.8%
                      SWE-bench68.3%
                      MMLU88.9%

                      Key Capabilities

                      • Instant code gen
                      • Repo-level understanding
                      • Test driven dev

                      Innovations

                        DeepSeek-V3.1

                        2025-08-21
                        Deep Seek
                        671B total·37B activeMoE

                        Incremental update with better instruction following.

                        GPQA81.4%
                        HLE25.1%
                        SWE-bench54.7%
                        MMLU87.6%

                        Key Capabilities

                        • Instruction following
                        • Long context
                        • Math improvements

                        Innovations

                          GPT-OSS

                          2025-08-05
                          OpenAI
                          Architecture not disclosed

                          Open source release of a highly capable model.

                          GPQA76.2%
                          HLE18.9%
                          SWE-bench45.1%
                          MMLU87.5%

                          Key Capabilities

                          • Open weights
                          • Community fine-tuning
                          • Transparency

                          Innovations

                            Claude Opus 4.1

                            2025-08-05
                            Anthropic
                            Architecture not disclosed

                            Massive reasoning model for complex R&D tasks.

                            GPQA88.5%
                            HLE38.9%
                            SWE-bench63.8%
                            MMLU91.5%
                            MMLU Pro87.92%

                            Key Capabilities

                            • Deep research
                            • Scientific discovery
                            • Long-horizon planning

                            Innovations

                              Genie 3

                              2025-08-05
                              Google
                              Architecture not disclosed

                              Google DeepMind's first real-time interactive general-purpose world model. From a text prompt it generates a navigable environment at 24 frames per second and 720p that stays consistent for a few minutes, and the world can be changed mid-simulation with further text prompts, which DeepMind calls promptable world events. Released as a limited research preview and positioned as a training environment for embodied agents rather than as a video generator; a public demo followed in January 2026 as Project Genie.

                              Key Capabilities

                              • Real-time interactive world generation
                              • Text-prompted world events during simulation
                              • Multi-minute visual consistency
                              • Agent training environments

                              Innovations

                              • First real-time interactive general-purpose world model
                              • Promptable world events
                              Jul

                              Qwen3-Coder

                              2025-07-15
                              Alibaba
                              Architecture not disclosed

                              Code-specialized model optimized for software engineering.

                              SWE-bench52.7%

                              Key Capabilities

                              • Code generation
                              • Agentic coding
                              • Multi-file editing
                              • Tool use

                              Innovations

                              • Code-specialized Training
                              • Agentic SWE Pipeline
                              • Multi-file Context

                              Grok 4

                              2025-07-09
                              X.AI
                              Architecture not disclosed

                              Next generation model with enhanced understanding of the physical world.

                              GPQA88.1%
                              HLE33.2%
                              SWE-bench60.4%
                              MMLU90.1%
                              MMLU Pro85.3%
                              ARC-AGI-215.9%

                              Key Capabilities

                              • Physical world model
                              • Video reasoning
                              • Robot control

                              Innovations

                                Jun

                                Gemini 2.5

                                2025-06-17
                                Google
                                Architecture not disclosed

                                Iterative update with enhanced speed and reasoning.

                                GPQA80.3%
                                HLE24.5%
                                SWE-bench55.4%
                                MMLU89.5%

                                Key Capabilities

                                • Flash/Pro variants
                                • Lower latency
                                • Higher accuracy

                                Innovations

                                • Nano Banana image generation

                                Seedance 1.0

                                2025-06-11
                                ByteDance
                                Architecture not disclosed

                                ByteDance's first Seedance video-generation model, introduced at the Volcano Engine Force conference. Fast text- and image-to-video generation with multi-shot narrative sequences.

                                Key Capabilities

                                • Text-to-video
                                • Image-to-video
                                • Multi-shot narratives
                                • Fast generation

                                Innovations

                                • Multi-shot narrative video generation
                                • First Seedance video model
                                Mistral
                                Architecture not disclosed

                                Mistral's first reasoning model. API and enterprise deployment only — no published weights. The 24B Magistral Small released the same day carries open weights and is tracked separately in the local catalog.

                                GPQA72.4%
                                HLE19.8%
                                SWE-bench48.9%
                                MMLU88.4%

                                Key Capabilities

                                • Enterprise focus
                                • Privacy first
                                • Specialized domains

                                Innovations

                                  May

                                  Claude Sonnet 4

                                  2025-05-22
                                  Anthropic
                                  Architecture not disclosed

                                  Next-generation Sonnet with improved instruction following and coding.

                                  GPQA86.4%
                                  HLE30.2%
                                  SWE-bench62.5%
                                  MMLU91.2%
                                  MMLU Pro83.86%

                                  Key Capabilities

                                  • Improved instruction following
                                  • Enhanced coding
                                  • Better reasoning

                                  Innovations

                                  • Improved Alignment
                                  • Better Calibration

                                  Claude Opus 4

                                  2025-05-22
                                  Anthropic
                                  Architecture not disclosed

                                  Frontier reasoning model for complex tasks and sustained autonomous work.

                                  GPQA88.5%
                                  HLE38.9%
                                  SWE-bench63.8%
                                  MMLU91.5%
                                  MMLU Pro86.17%

                                  Key Capabilities

                                  • Deep research
                                  • Sustained autonomy
                                  • Complex reasoning

                                  Innovations

                                  • Extended Autonomy
                                  • Deep Analysis
                                  • Multi-hour Tasks

                                  Veo 3

                                  2025-05-13
                                  Google
                                  Architecture not disclosed

                                  Major update with synchronized audio generation.

                                  Key Capabilities

                                  • Synchronized audio
                                  • Video generation
                                  • Cinematic controls

                                  Innovations

                                  • Native Audio Generation
                                  • Flow Platform
                                  • Gemini Integration
                                  Apr

                                  Qwen3

                                  2025-04-29
                                  Alibaba
                                  235B total·22B activeMoE

                                  235B MoE flagship with hybrid thinking modes.

                                  GPQA68.9%
                                  SWE-bench42%
                                  MMLU89.5%

                                  Key Capabilities

                                  • 235B MoE
                                  • Hybrid thinking
                                  • 119 languages
                                  • 128k context

                                  Innovations

                                  • Thinking/Non-thinking Modes
                                  • MoE Architecture
                                  • Agentic Coding

                                  o3

                                  2025-04-16
                                  OpenAI
                                  Architecture not disclosed

                                  Full reasoning model with significant improvements over o1.

                                  GPQA82.5%
                                  HLE28.4%
                                  SWE-bench69.1%
                                  MMLU91.6%
                                  MMLU Pro85.59%
                                  ARC-AGI-175.7%

                                  Key Capabilities

                                  • Advanced reasoning
                                  • Tool use
                                  • Agentic coding

                                  Innovations

                                  • Improved Reasoning Chain
                                  • Tool Integration
                                  • Multi-step Planning

                                  o4-mini

                                  2025-04-16
                                  OpenAI
                                  Architecture not disclosed

                                  Fast, cost-effective reasoning model released alongside o3.

                                  GPQA76.8%
                                  SWE-bench68.1%
                                  MMLU89.5%

                                  Key Capabilities

                                  • Fast reasoning
                                  • Cost-effective
                                  • Tool use

                                  Innovations

                                    Seedream 3.0

                                    2025-04-16
                                    ByteDance
                                    Architecture not disclosed

                                    ByteDance's bilingual (Chinese/English) text-to-image foundation model. Native 2048×2048 output with fast inference (~3s for a 1K image on an A100). Available via the Doubao and Jimeng apps.

                                    Key Capabilities

                                    • Text-to-Image
                                    • Bilingual (Chinese/English) prompts
                                    • Native 2048×2048 output
                                    • Fast inference (~3s at 1K)

                                    Innovations

                                    • Bilingual text-to-image foundation model
                                    • Native 2K output

                                    Kling 2.0

                                    2025-04-15
                                    Kuaishou
                                    Architecture not disclosed

                                    Cinematic quality video with advanced physics and multi-element editing.

                                    Key Capabilities

                                    • Cinematic 1080p
                                    • Physics simulation
                                    • Multi-element editor
                                    • Video generation

                                    Innovations

                                    • Multi-Modal Visual Language
                                    • Multi-Element Editor
                                    • Advanced Physics

                                    Llama 4 Scout

                                    2025-04-05
                                    Meta
                                    109B total·17B activeMoE

                                    109B MoE (17B active) multimodal model with 10M token context window. Outperforms Gemma 3, Gemini 2.0 Flash, and Mistral 3.1 Small.

                                    GPQA57.2%
                                    MMLU85.8%

                                    Key Capabilities

                                    • 109B total / 17B active MoE
                                    • 10M context window
                                    • Natively multimodal
                                    • 16 experts, 1 active
                                    • Open weights (Llama license)

                                    Innovations

                                    • Early Fusion Architecture
                                    • 10M Token Context (Longest Open-Weight)
                                    • Interleaved Attention / MoE Blocks
                                    Meta
                                    400B total·17B activeMoE

                                    400B MoE (17B active) multimodal model. Matches Gemini 2.5 Pro and DeepSeek-V3 on reasoning benchmarks at significantly lower serving cost.

                                    GPQA69.8%
                                    MMLU92.4%

                                    Key Capabilities

                                    • 400B total / 17B active MoE
                                    • 128 experts, 1 active
                                    • Natively multimodal
                                    • 1M context window
                                    • Open weights (Llama license)

                                    Innovations

                                    • 128-Expert MoE (Largest Open-Weight MoE)
                                    • Early Fusion Architecture
                                    • Frontier Performance at Low Active Params
                                    Meta
                                    2T total·288B activeMoE

                                    2T MoE (288B active) teacher model. Announced alongside Scout/Maverick, still in training at launch. Tops benchmarks in STEM reasoning.

                                    Key Capabilities

                                    • 2T total / 288B active MoE
                                    • 16 experts, 4 active
                                    • STEM reasoning
                                    • Natively multimodal
                                    • In training at announcement

                                    Innovations

                                    • Largest Llama Model Ever
                                    • Teacher Model Architecture
                                    • 288B Active Parameters

                                    Midjourney V7

                                    2025-04-03
                                    Midjourney
                                    Architecture not disclosed

                                    Stunning precision with improved text and image handling.

                                    Key Capabilities

                                    • Draft Mode
                                    • Omni Reference
                                    • Rich textures
                                    • Better hands/bodies

                                    Innovations

                                    • Draft Mode
                                    • Omni Reference
                                    • Enhanced Precision

                                    Runway Gen-4

                                    2025-04-01
                                    Runway
                                    Architecture not disclosed

                                    Next-gen model with improved character and environment consistency.

                                    Key Capabilities

                                    • Character consistency
                                    • Environment persistence
                                    • 16s videos
                                    • Gen-4 Turbo

                                    Innovations

                                    • Consistency Engine
                                    • Extended Duration
                                    • Turbo Mode
                                    Mar

                                    Gemini 2.5 Pro

                                    2025-03-25
                                    Google
                                    Architecture not disclosed

                                    Thinking model with enhanced reasoning capabilities.

                                    GPQA80.3%
                                    HLE24.5%
                                    SWE-bench55.4%
                                    MMLU89.5%
                                    MMLU Pro84.06%

                                    Key Capabilities

                                    • Thinking/reasoning
                                    • Code generation
                                    • Multimodal

                                    Innovations

                                    • Thinking Mode
                                    • Enhanced Reasoning
                                    • Experimental Preview

                                    Hunyuan-T1

                                    2025-03-21
                                    Tencent

                                    Tencent's deep-thinking reasoning flagship, built on the TurboS Hybrid-Transformer-Mamba MoE base and post-trained with large-scale curriculum reinforcement learning. First Hunyuan entry in the frontier reasoning race.

                                    GPQA69.3%
                                    MMLU Pro87.2%

                                    Key Capabilities

                                    • Deep-thinking chain-of-thought reasoning
                                    • Hybrid-Transformer-Mamba MoE base (TurboS)
                                    • Fast decoding on long inputs

                                    Innovations

                                    • First ultra-large Mamba-hybrid reasoning model
                                    • Curriculum reinforcement-learning post-training

                                    Command A

                                    2025-03-13
                                    Cohere
                                    111B

                                    111B Dense Enterprise-Flagship (command-a-03-2025), 256k Kontext, 23 Sprachen, Fokus auf Tool-Use, agentische Workflows und Translation. Läuft auf wenigen GPUs, offene Gewichte. Direkter Vorgänger von Command A+. Vectara HHEM 9.3% Halluzination.

                                    GPQA52.7%
                                    MMLU Pro71.2%

                                    Key Capabilities

                                    • 111B Dense
                                    • Enterprise / Tool-Use / Agentic
                                    • 23 Sprachen
                                    • 256k Context
                                    • Open weights

                                    Innovations

                                    • Effizientes Enterprise-Flagship
                                    • Open weights
                                    Feb

                                    ERNIE X1

                                    2025-02-28
                                    Baidu
                                    Architecture not disclosed

                                    Reasoning model competing with DeepSeek-R1.

                                    Key Capabilities

                                    • Deep reasoning
                                    • Chain-of-thought
                                    • Math & science

                                    Innovations

                                    • Deliberative Reasoning
                                    • DeepSeek-R1 Competitor
                                    • Thinking Tokens

                                    GPT-4.5

                                    2025-02-27
                                    OpenAI
                                    Architecture not disclosed

                                    Largest and most capable GPT model with reduced hallucinations and improved EQ.

                                    GPQA71.4%
                                    SWE-bench38.3%
                                    MMLU90.2%

                                    Key Capabilities

                                    • Reduced hallucinations
                                    • Improved EQ
                                    • Larger context

                                    Innovations

                                    • Unsupervised Learning Scale
                                    • Improved Calibration
                                    Anthropic
                                    Architecture not disclosed

                                    Hybrid model with extended thinking for complex reasoning tasks.

                                    GPQA84.8%
                                    HLE26.7%
                                    SWE-bench56.2%
                                    MMLU90.7%
                                    MMLU Pro82.73%

                                    Key Capabilities

                                    • Extended thinking
                                    • Enhanced coding
                                    • Agentic reliability

                                    Innovations

                                    • Extended Thinking Mode
                                    • Hybrid Reasoning
                                    • Improved Tool Use

                                    Grok-3

                                    2025-02-17
                                    X.AI
                                    Architecture not disclosed

                                    Massive scale model trained on the world's largest supercomputer.

                                    GPQA81.2%
                                    HLE25.6%
                                    SWE-bench51.8%
                                    MMLU86.5%

                                    Key Capabilities

                                    • 100k H100 training
                                    • Super-intelligence
                                    • Multimodal

                                    Innovations

                                      Grok-3 Mini

                                      2025-02-17
                                      X.AI
                                      Architecture not disclosed

                                      Lightweight reasoning model from xAI. Faster and cheaper than Grok-3 with configurable thinking for speed/quality tradeoffs.

                                      Key Capabilities

                                      • Lightweight reasoning
                                      • Configurable thinking mode
                                      • Faster inference than Grok-3
                                      • Cost-optimized

                                      Innovations

                                      • Thinking Mode (Low/High)
                                      • Reasoning at Mini Scale
                                      Google
                                      Architecture not disclosed

                                      Next-gen Flash model with native tool use and multimodal generation.

                                      GPQA62.1%
                                      MMLU83.8%

                                      Key Capabilities

                                      • Native tool use
                                      • Multimodal generation
                                      • Fast inference

                                      Innovations

                                      • Native Tool Use
                                      • Multimodal Output
                                      • Agentic Capabilities
                                      Jan

                                      o3-mini

                                      2025-01-31
                                      OpenAI
                                      Architecture not disclosed

                                      Smaller, faster reasoning model bringing o-series capabilities to a more accessible form.

                                      GPQA79.7%
                                      HLE13.4%
                                      SWE-bench49.3%
                                      MMLU86.7%

                                      Key Capabilities

                                      • Reasoning
                                      • Fast inference
                                      • Cost-effective

                                      Innovations

                                      • Compact Reasoning
                                      • Efficient CoT
                                      • Accessible Intelligence

                                      Mistral Small 3

                                      2025-01-30
                                      Mistral
                                      24B

                                      24B open-weight model under Apache 2.0 with strong efficiency.

                                      MMLU81%

                                      Key Capabilities

                                      • 24B Parameters
                                      • Apache 2.0
                                      • On-device capable

                                      Innovations

                                      • Efficient 24B Scale
                                      • Apache 2.0 License
                                      • Edge Deployment

                                      Operator

                                      2025-01-23
                                      OpenAI
                                      Architecture not disclosed

                                      Agentic system capable of executing complex multi-step tasks on a computer.

                                      HLE22.1%
                                      SWE-bench52.3%

                                      Key Capabilities

                                      • Computer use
                                      • Autonomous browsing
                                      • Task execution

                                      Innovations

                                      • Computer Use
                                      • Autonomous Agents
                                      • Multi-step Task Execution

                                      DeepSeek-R1

                                      2025-01-20
                                      Deep Seek
                                      671B total·37B activeMoE

                                      Reasoning model capable of self-verification.

                                      GPQA71.5%
                                      HLE18.4%
                                      SWE-bench49.2%
                                      MMLU90.8%
                                      MMLU Pro83.18%

                                      Key Capabilities

                                      • Chain of thought
                                      • Self-verification
                                      • RL training

                                      Innovations

                                      • Pure RL Reasoning
                                      • Self-verification
                                      • Emergent CoT Behavior
                                      2024
                                      Dec

                                      DeepSeek-V3

                                      2024-12-25
                                      Deep Seek
                                      671B total·37B activeMoE

                                      Latest iteration with improved benchmarks.

                                      GPQA59.1%
                                      HLE11.8%
                                      SWE-bench42%
                                      MMLU88.5%

                                      Key Capabilities

                                      • SOTA Open Weights
                                      • FP8 Training
                                      • Multi-token prediction

                                      Innovations

                                      • FP8 Training
                                      • Multi-token Prediction
                                      • Auxiliary Loss-Free Load Balancing

                                      ERNIE 4.5

                                      2024-12-20
                                      Baidu
                                      Architecture not disclosed

                                      Improved reasoning and multimodal capabilities.

                                      MMLU82%

                                      Key Capabilities

                                      • Enhanced reasoning
                                      • Multimodal
                                      • Long context

                                      Innovations

                                      • Deeper Reasoning
                                      • Improved Alignment
                                      • Extended Context

                                      Kling 1.6

                                      2024-12-19
                                      Kuaishou
                                      Architecture not disclosed

                                      Enhanced realism with better motion and prompt understanding.

                                      Key Capabilities

                                      • Enhanced realism
                                      • Better expressions
                                      • Video generation
                                      • Style preservation

                                      Innovations

                                      • DeepSeek Prompt Understanding
                                      • Elements Feature
                                      • 195% Quality Boost

                                      Veo 2

                                      2024-12-16
                                      Google
                                      Architecture not disclosed

                                      Enhanced video model with 4K resolution support.

                                      Key Capabilities

                                      • 4K resolution
                                      • Better physics
                                      • Improved realism

                                      Innovations

                                      • 4K Video
                                      • Enhanced Physics
                                      • VideoFX Integration

                                      Llama 3.3

                                      2024-12-06
                                      Meta
                                      70B

                                      Efficient 70B text model matching Llama 3.1 405B on key benchmarks.

                                      GPQA46.7%
                                      MMLU86.3%

                                      Key Capabilities

                                      • 70B Parameters
                                      • 405B-level performance
                                      • Cost efficient

                                      Innovations

                                      • Efficiency Breakthrough
                                      • Distillation Techniques
                                      Nov

                                      QwQ-32B-Preview

                                      2024-11-27
                                      Alibaba
                                      32B

                                      Reasoning-focused model with chain-of-thought capabilities.

                                      GPQA54.5%

                                      Key Capabilities

                                      • 32B Parameters
                                      • Chain-of-thought
                                      • Math reasoning
                                      • Self-reflection

                                      Innovations

                                      • Deliberative Reasoning
                                      • Extended Thinking
                                      • Self-Correction

                                      Flux.1 Tools

                                      2024-11-21
                                      Black Forest Labs
                                      Architecture not disclosed

                                      Suite of editing tools including Fill, Depth, Canny, and Redux.

                                      Key Capabilities

                                      • Inpainting
                                      • Outpainting
                                      • Depth control
                                      • Edge control

                                      Innovations

                                      • Flux.1 Fill
                                      • Flux.1 Depth
                                      • Flux.1 Canny
                                      • Flux.1 Redux

                                      Pixtral Large

                                      2024-11-18
                                      Mistral
                                      124B

                                      124B multimodal model with frontier-class vision capabilities.

                                      Key Capabilities

                                      • 124B Parameters
                                      • Frontier vision
                                      • Multimodal

                                      Innovations

                                      • Large-scale Vision
                                      • 128k Context
                                      • Multimodal Reasoning
                                      Oct
                                      Stability AI
                                      2.5B

                                      Compact 2.5B parameter variant optimized for consumer GPUs. Runs on hardware with 4-16 GB VRAM.

                                      Key Capabilities

                                      • 2.5B parameters
                                      • Text-to-image
                                      • Consumer GPU optimized (4-16 GB VRAM)
                                      • Open weights (Stability Community License)

                                      Innovations

                                      • Efficient MMDiT at 2.5B Scale
                                      • Consumer Hardware Target

                                      Aya Expanse 32B

                                      2024-10-24
                                      Cohere
                                      32B

                                      32B multilinguales Open-Weight-Modell aus der Aya-Forschung von Cohere Labs, 23 Sprachen, Schwerpunkt auf non-English-Evaluierungen (Global-MMLU, mArenaHard). Die englische Standard-Suite reportet der Vendor nicht, multilinguale Stärke und offene Gewichte sind der Fokus. Vectara HHEM 10.9% Halluzination.

                                      Key Capabilities

                                      • 32B Dense
                                      • Multilingual (23 Sprachen)
                                      • Open weights
                                      • Aya-Forschungslinie

                                      Innovations

                                      • Multilingual-fokussiertes Open-Weight-Modell
                                      • Data Arbitrage und multilinguales Preference-Training
                                      Anthropic
                                      Architecture not disclosed

                                      Upgraded Sonnet with computer use capabilities.

                                      GPQA65%
                                      HLE11.5%
                                      SWE-bench49%
                                      MMLU88.3%

                                      Key Capabilities

                                      • Computer use
                                      • Enhanced coding
                                      • Tool use

                                      Innovations

                                      • Computer Use
                                      • Agent Capabilities
                                      • Real-world Interaction
                                      Anthropic
                                      Architecture not disclosed

                                      Fastest model in its intelligence class.

                                      GPQA61.8%
                                      HLE10.1%
                                      SWE-bench28.7%
                                      MMLU85.9%

                                      Key Capabilities

                                      • Low latency
                                      • High intelligence
                                      • Tool use

                                      Innovations

                                        Stability AI
                                        8.1B

                                        Flagship 8.1B parameter open-weight image model with improved quality and prompt adherence. Available in standard and Turbo variants.

                                        Key Capabilities

                                        • 8.1B parameters
                                        • Text-to-image
                                        • Large + Large Turbo variants
                                        • Open weights (Stability Community License)

                                        Innovations

                                        • MMDiT-X Architecture
                                        • QK-Normalization
                                        • Turbo Distillation Variant

                                        Flux 1.1 Pro

                                        2024-10-02
                                        Black Forest Labs
                                        Architecture not disclosed

                                        Improved flagship model with enhanced quality and speed.

                                        Key Capabilities

                                        • Faster generation
                                        • Better quality
                                        • Professional use

                                        Innovations

                                        • Speed Improvements
                                        • Quality Enhancement
                                        • Pro Features
                                        Sep

                                        Qwen2.5

                                        2024-09-19
                                        Alibaba
                                        72B

                                        Flagship 72B model rivaling GPT-4o and Claude 3.5 on key benchmarks.

                                        GPQA49%
                                        SWE-bench23%
                                        MMLU88.5%

                                        Key Capabilities

                                        • 72B Parameters
                                        • 128k context
                                        • Structured output
                                        • Tool use

                                        Innovations

                                        • Post-training Scaling
                                        • Improved Math & Code
                                        • Long-context Generation

                                        Kling 1.5

                                        2024-09-19
                                        Kuaishou
                                        Architecture not disclosed

                                        Improved quality with Motion Brush and camera controls.

                                        Key Capabilities

                                        • 1080p HD
                                        • Motion Brush
                                        • Camera controls
                                        • Video generation

                                        Innovations

                                        • Motion Brush
                                        • Camera Movement Controls
                                        • 95% Quality Improvement

                                        Pixtral 12B

                                        2024-09-17
                                        Mistral
                                        12B

                                        First multimodal model from Mistral with vision capabilities.

                                        Key Capabilities

                                        • Vision understanding
                                        • 12B Parameters
                                        • Open weights

                                        Innovations

                                        • Multimodal Vision
                                        • Variable Image Resolution
                                        • Native Vision Encoder

                                        OpenAI o1

                                        2024-09-12
                                        OpenAI
                                        Architecture not disclosed

                                        Reasoning model designed to spend more time thinking before responding.

                                        GPQA78%
                                        HLE15.2%
                                        SWE-bench48.9%
                                        MMLU91.8%
                                        MMLU Pro83.49%

                                        Key Capabilities

                                        • Chain of thought
                                        • Complex math/coding
                                        • Self-correction

                                        Innovations

                                        • Chain of Thought Reasoning
                                        • Test-time Compute Scaling
                                        • Self-correction via RL
                                        Aug

                                        Imagen 3

                                        2024-08-28
                                        Google
                                        Architecture not disclosed

                                        Highest quality text-to-image model from Google DeepMind.

                                        Key Capabilities

                                        • Text-to-Image
                                        • Photorealism
                                        • Text rendering
                                        • Diverse styles

                                        Innovations

                                        • Enhanced Photorealism
                                        • Text Rendering Quality
                                        • SynthID Watermarking

                                        Grok-2

                                        2024-08-14
                                        X.AI
                                        Architecture not disclosed

                                        State-of-the-art language model with image generation.

                                        GPQA58.2%
                                        HLE10.5%
                                        SWE-bench24.7%
                                        MMLU85.6%

                                        Key Capabilities

                                        • FLUX.1 Image Gen
                                        • SOTA Reasoning
                                        • Chat interface

                                        Innovations

                                        • FLUX.1 Integration
                                        • Native Image Generation
                                        • Unrestricted Content Policy

                                        FLUX.1 [dev]

                                        2024-08-01
                                        Black Forest Labs
                                        12B

                                        Open-weight text-to-image model from Black Forest Labs. Guidance-distilled 12B param model for non-commercial use.

                                        Key Capabilities

                                        • 12B parameters
                                        • Text-to-image
                                        • Guidance distilled
                                        • Open weights (non-commercial)

                                        Innovations

                                        • Flow Matching Architecture
                                        • Hybrid Transformer (MMDiT)
                                        • Guidance Distillation
                                        Black Forest Labs
                                        12B

                                        Fastest FLUX variant, Apache 2.0 licensed. Optimized for local inference and rapid prototyping.

                                        Key Capabilities

                                        • 12B parameters
                                        • Text-to-image
                                        • Fastest FLUX variant
                                        • Apache 2.0 license

                                        Innovations

                                        • Timestep Distillation (4 steps)
                                        • Apache 2.0 Open Weight Image Gen
                                        Jul

                                        Midjourney V6.1

                                        2024-07-30
                                        Midjourney
                                        Architecture not disclosed

                                        Faster generation with more coherent images and precise details.

                                        Key Capabilities

                                        • 25% faster
                                        • Coherent images
                                        • Precise textures

                                        Innovations

                                        • Speed Optimization
                                        • Detail Enhancement
                                        • Texture Precision

                                        Mistral Large 2

                                        2024-07-24
                                        Mistral
                                        123B

                                        Significant improvements in code generation and reasoning.

                                        GPQA55.1%
                                        HLE8.9%
                                        SWE-bench32.6%
                                        MMLU84%

                                        Key Capabilities

                                        • 123B Parameters
                                        • Code generation
                                        • Function calling

                                        Innovations

                                        • 123B Parameters
                                        • Code-focused Training
                                        • Advanced Function Calling

                                        Llama 3.1

                                        2024-07-23
                                        Meta
                                        405B

                                        Massive 405B open-source model competing with frontier closed models.

                                        GPQA46.7%
                                        MMLU88.6%

                                        Key Capabilities

                                        • 405B Parameters
                                        • 128k Context
                                        • Tool use
                                        • Multilingual

                                        Innovations

                                        • 405B Open Weights
                                        • 128k Context
                                        • Synthetic Data Scaling

                                        Mistral Nemo

                                        2024-07-18
                                        Mistral
                                        12B

                                        12B parameter model developed in collaboration with NVIDIA.

                                        MMLU68%

                                        Key Capabilities

                                        • 12B Parameters
                                        • Apache 2.0
                                        • 128k context

                                        Innovations

                                        • NVIDIA Collaboration
                                        • Tekken Tokenizer
                                        • Efficient 12B Scale
                                        Jun

                                        ERNIE 4.0 Turbo

                                        2024-06-28
                                        Baidu
                                        Architecture not disclosed

                                        Speed-optimized variant with reduced latency.

                                        MMLU78.5%

                                        Key Capabilities

                                        • Faster inference
                                        • Advanced reasoning
                                        • Tool use

                                        Innovations

                                        • Optimized Inference
                                        • Reduced Latency
                                        • Cost Efficiency
                                        Anthropic
                                        Architecture not disclosed

                                        Significant leap in coding and reasoning capabilities.

                                        GPQA59.4%
                                        HLE12.4%
                                        SWE-bench49%
                                        MMLU88.7%

                                        Key Capabilities

                                        • SOTA Coding
                                        • Artifacts UI
                                        • Fast inference

                                        Innovations

                                        • Artifacts UI
                                        • Real-time Collaboration
                                        • Interactive Content Generation
                                        Runway
                                        Architecture not disclosed

                                        Advanced video generation with creative control and motion tools.

                                        Key Capabilities

                                        • Creative control
                                        • Motion brush
                                        • Camera controls
                                        • Director mode

                                        Innovations

                                        • Advanced Camera Controls
                                        • Motion Brush
                                        • Keyframing
                                        Stability AI
                                        Architecture not disclosed

                                        Next-gen architecture with Multimodal Diffusion Transformer.

                                        Key Capabilities

                                        • MMDiT architecture
                                        • Improved text rendering
                                        • Better composition

                                        Innovations

                                        • Multimodal Diffusion Transformer
                                        • Flow Matching
                                        • Triple Text Encoders

                                        Qwen2

                                        2024-06-07
                                        Alibaba
                                        72B

                                        Major upgrade with 72B flagship, competitive with leading Western models.

                                        GPQA30.4%
                                        MMLU84.2%

                                        Key Capabilities

                                        • 72B Parameters
                                        • 128k context
                                        • Multilingual (29 languages)

                                        Innovations

                                        • GQA Architecture
                                        • Dual Chunk Attention
                                        • YARN Scaling

                                        Kling 1.0

                                        2024-06-06
                                        Kuaishou
                                        Architecture not disclosed

                                        First consumer-accessible video generation model with up to 2 min output.

                                        Key Capabilities

                                        • Text-to-video
                                        • 1080p/30fps
                                        • Up to 2 min
                                        • Complex motion

                                        Innovations

                                        • DiT Architecture
                                        • 3D VAE Compression
                                        • Spatiotemporal Simulation
                                        May
                                        Google
                                        Architecture not disclosed

                                        Lightweight model optimized for speed and efficiency.

                                        GPQA51.2%
                                        HLE7.8%
                                        MMLU78.9%

                                        Key Capabilities

                                        • Low latency
                                        • High throughput
                                        • Cost effective

                                        Innovations

                                        • Efficient MoE
                                        • Low Latency Optimization
                                        • Cost-effective Inference

                                        Veo

                                        2024-05-14
                                        Google
                                        Architecture not disclosed

                                        Google's first text-to-video model with high-quality video generation.

                                        Key Capabilities

                                        • Text-to-video
                                        • High quality
                                        • 1080p output

                                        Innovations

                                        • Video Generation
                                        • Temporal Consistency
                                        • Cinematic Quality

                                        GPT-4o

                                        2024-05-13
                                        OpenAI
                                        Architecture not disclosed

                                        Omni model accepting text, audio, and image inputs.

                                        GPQA53.6%
                                        HLE8.5%
                                        SWE-bench38%
                                        MMLU88.7%

                                        Key Capabilities

                                        • Real-time voice
                                        • Native multimodal
                                        • Faster inference

                                        Innovations

                                        • Native Omni-modal
                                        • Real-time Voice (320ms)
                                        • End-to-end Multimodal Training

                                        DeepSeek-V2

                                        2024-05-01
                                        Deep Seek
                                        236B total·21B activeMoE

                                        Strong performance at a lower cost.

                                        GPQA48.3%
                                        HLE6.1%
                                        SWE-bench18.9%
                                        MMLU77.8%

                                        Key Capabilities

                                        • MLA Architecture
                                        • GPT-4 class
                                        • Extremely cheap API

                                        Innovations

                                        • Multi-head Latent Attention (MLA)
                                        • KV Cache Compression (93.3%)
                                        • DeepSeekMoE Architecture
                                        Apr

                                        Command R+

                                        2024-04-24
                                        Cohere
                                        104B

                                        104B Dense, RAG- und Tool-Use-optimiertes Open-Weight-Flagship (CC-BY-NC), 128k Kontext, citation-fähiges Retrieval. Starkes Grounding beim Zusammenfassen (Vectara HHEM 6.9% Halluzination für den 08-2024-Build), akademische Benchmarks dagegen schwach. Direkter Vorläufer der Command-A-Linie.

                                        GPQA32.3%
                                        MMLU Pro43.2%

                                        Key Capabilities

                                        • 104B Dense
                                        • RAG-optimiert
                                        • Tool-Use
                                        • Multilingual (10 Sprachen)
                                        • 128k Context

                                        Innovations

                                        • Citation-fähiges RAG
                                        • Open weights (CC-BY-NC)

                                        Llama 3

                                        2024-04-18
                                        Meta
                                        Architecture not disclosed

                                        Next generation open-source LLM with 8B and 70B variants.

                                        GPQA32.8%
                                        MMLU82%

                                        Key Capabilities

                                        • 8B/70B sizes
                                        • Improved reasoning
                                        • Enhanced coding

                                        Innovations

                                        • Improved Tokenizer
                                        • GQA Across All Sizes
                                        • Llama Guard Safety
                                        Mar

                                        Grok-1.5

                                        2024-03-29
                                        X.AI
                                        Architecture not disclosed

                                        Improved reasoning and coding capabilities.

                                        Key Capabilities

                                        • 128k Context
                                        • Math reasoning
                                        • Coding benchmarks

                                        Innovations

                                        • 128k Context Window
                                        • Enhanced Math Reasoning
                                        • Coding Benchmarks

                                        Grok-1

                                        2024-03-17
                                        X.AI
                                        314B total·86B activeMoE

                                        314B MoE (86B active) open-weight model released under Apache 2.0. Largest open-weight model at the time of release.

                                        Key Capabilities

                                        • 314B total / 86B active MoE
                                        • 8 experts, 2 active
                                        • Apache 2.0 license
                                        • Largest open-weight model at release

                                        Innovations

                                        • Largest Open-Weight MoE at Launch
                                        • Apache 2.0 Full Weight Release

                                        Claude 3 Opus

                                        2024-03-04
                                        Anthropic
                                        Architecture not disclosed

                                        Haiku, Sonnet, and Opus models.

                                        GPQA50.4%
                                        HLE6.5%
                                        SWE-bench13.9%
                                        MMLU86.8%

                                        Key Capabilities

                                        • Vision capabilities
                                        • Near-human nuance
                                        • Opus reasoning

                                        Innovations

                                        • Vision Capabilities
                                        • Three-tier Model Family
                                        • Near-human Reasoning
                                        Feb

                                        Mistral Large

                                        2024-02-26
                                        Mistral
                                        Architecture not disclosed

                                        Flagship model with top-tier reasoning capabilities.

                                        GPQA42.5%
                                        HLE5.2%
                                        SWE-bench18.3%
                                        MMLU81.2%

                                        Key Capabilities

                                        • 32k Context
                                        • Multi-lingual
                                        • Strong reasoning

                                        Innovations

                                        • 32k Context Window
                                        • Multi-lingual Excellence
                                        • Function Calling

                                        Sora

                                        2024-02-15
                                        OpenAI
                                        Architecture not disclosed

                                        Text-to-video model capable of creating realistic and imaginative scenes.

                                        Key Capabilities

                                        • Text-to-Video
                                        • Physics simulation
                                        • High fidelity

                                        Innovations

                                        • World Simulation
                                        • Diffusion Transformers
                                        • Temporal Coherence

                                        Gemini 1.5 Pro

                                        2024-02-15
                                        Google
                                        Architecture not disclosed

                                        Mid-size multimodal model optimized for scaling.

                                        GPQA56.8%
                                        HLE9.1%
                                        SWE-bench40.5%
                                        MMLU85.9%

                                        Key Capabilities

                                        • 1M+ Context window
                                        • Video understanding
                                        • In-context learning

                                        Innovations

                                        • 1M+ Context Window
                                        • MoE Architecture
                                        • Video Understanding

                                        Qwen 1.5

                                        2024-02-04
                                        Alibaba
                                        Architecture not disclosed

                                        Multi-size release from 0.5B to 110B with strong multilingual support.

                                        MMLU75.7%

                                        Key Capabilities

                                        • 0.5B–110B sizes
                                        • Multilingual
                                        • 32k context

                                        Innovations

                                        • Scalable Model Family
                                        • RLHF Alignment
                                        • Extended Context
                                        Jan

                                        DeepSeek-MoE

                                        2024-01-01
                                        Deep Seek
                                        14.6B total·2.8B activeMoE

                                        Mixture-of-Experts architecture.

                                        Key Capabilities

                                        • MoE efficiency
                                        • High throughput
                                        • Cost effective

                                        Innovations

                                        • MoE Efficiency
                                        • Cost-effective Training
                                        • High Throughput
                                        2023
                                        Dec

                                        Midjourney V6

                                        2023-12-21
                                        Midjourney
                                        Architecture not disclosed

                                        Enhanced prompt accuracy and image coherence.

                                        Key Capabilities

                                        • Better prompt accuracy
                                        • Improved coherence
                                        • Advanced image prompting

                                        Innovations

                                        • Prompt Accuracy
                                        • Image Coherence
                                        • Advanced Prompting

                                        Mixtral 8x7B

                                        2023-12-11
                                        Mistral
                                        46.7B total·12.9B activeMoE

                                        Sparse Mixture-of-Experts model.

                                        Key Capabilities

                                        • MoE Architecture
                                        • GPT-3.5 level
                                        • Open weights

                                        Innovations

                                        • Sparse MoE Architecture
                                        • Router Network
                                        • Expert Selection Mechanism

                                        Gemini 1.0

                                        2023-12-06
                                        Google
                                        Architecture not disclosed

                                        Multimodal model built from the ground up.

                                        GPQA32.1%
                                        HLE3.5%
                                        MMLU90%

                                        Key Capabilities

                                        • Native multimodal
                                        • Nano/Pro/Ultra sizes
                                        • AlphaCode 2 integration

                                        Innovations

                                        • Native Multimodality
                                        • Ground-up Multimodal Training
                                        • AlphaCode 2 Integration
                                        Nov

                                        GPT-4 Turbo

                                        2023-11-06
                                        OpenAI
                                        Architecture not disclosed

                                        128k context window with knowledge up to April 2023, announced at DevDay.

                                        GPQA38.2%
                                        HLE5.1%
                                        SWE-bench4.2%
                                        MMLU86.4%

                                        Key Capabilities

                                        • 128k Context
                                        • JSON mode
                                        • Cheaper pricing

                                        Innovations

                                        • 128k Context Window
                                        • JSON Mode
                                        • Reproducible Outputs

                                        DeepSeek-Coder

                                        2023-11-02
                                        Deep Seek
                                        Architecture not disclosed

                                        Open-source code generation model.

                                        Key Capabilities

                                        • Code specialization
                                        • Open source
                                        • Various sizes

                                        Innovations

                                        • Code Specialization
                                        • Multiple Model Sizes
                                        • Open Source Focus
                                        Oct

                                        ERNIE 4.0

                                        2023-10-17
                                        Baidu
                                        Architecture not disclosed

                                        Major upgrade rivaling GPT-4 on Chinese language tasks.

                                        MMLU76%

                                        Key Capabilities

                                        • Advanced reasoning
                                        • Code generation
                                        • Multimodal

                                        Innovations

                                        • Multi-step Reasoning
                                        • Chinese Benchmark Leader
                                        • Plugin Ecosystem

                                        DALL-E 3

                                        2023-10-03
                                        OpenAI
                                        Architecture not disclosed

                                        Major image generation upgrade integrated directly into ChatGPT.

                                        Key Capabilities

                                        • ChatGPT integration
                                        • Improved text rendering
                                        • Higher fidelity

                                        Innovations

                                        • ChatGPT Integration
                                        • Prompt Rewriting
                                        • Text Rendering in Images
                                        Sep

                                        Mistral 7B

                                        2023-09-27
                                        Mistral
                                        Architecture not disclosed

                                        Powerful open-weight model.

                                        Key Capabilities

                                        • Open weights
                                        • Efficient attention
                                        • High performance/size

                                        Innovations

                                        • Grouped Query Attention
                                        • Sliding Window Attention
                                        • Efficient Open Weights
                                        Aug

                                        Code Llama

                                        2023-08-24
                                        Meta
                                        Architecture not disclosed

                                        Code-specialized Llama model for code generation and understanding.

                                        Key Capabilities

                                        • Code generation
                                        • Infilling
                                        • Instruction following

                                        Innovations

                                        • Code Specialization
                                        • Infilling Capability
                                        • Long Context Fine-tuning

                                        Qwen-7B

                                        2023-08-03
                                        Alibaba
                                        7B

                                        Alibaba's first open-source large language model.

                                        Key Capabilities

                                        • 7B Parameters
                                        • Multilingual
                                        • Code generation

                                        Innovations

                                        • Open-source Chinese LLM
                                        • Multilingual Training
                                        • Tool Use
                                        Jul
                                        Stability AI
                                        Architecture not disclosed

                                        Major upgrade with native 1024x1024 resolution and improved generation.

                                        Key Capabilities

                                        • 1024x1024 native
                                        • Two-stage pipeline
                                        • Improved aesthetics

                                        Innovations

                                        • Higher Resolution
                                        • Refiner Model
                                        • Enhanced Prompt Understanding

                                        Llama 2

                                        2023-07-18
                                        Meta
                                        Architecture not disclosed

                                        Open-source LLM available for commercial use.

                                        MMLU68.9%

                                        Key Capabilities

                                        • 7B/13B/70B sizes
                                        • Commercial license
                                        • Chat fine-tuning

                                        Innovations

                                        • Open Commercial License
                                        • RLHF Chat Models
                                        • Safety Red-teaming

                                        Claude 2

                                        2023-07-11
                                        Anthropic
                                        Architecture not disclosed

                                        Improved performance and longer context window.

                                        GPQA29.5%
                                        HLE2.8%
                                        SWE-bench4.8%
                                        MMLU78.5%

                                        Key Capabilities

                                        • 100k Context
                                        • PDF analysis
                                        • Improved coding

                                        Innovations

                                        • 100k Context Window
                                        • Long Document Analysis
                                        • Extended Context Handling
                                        Mar

                                        Bard

                                        2023-03-21
                                        Google
                                        Architecture not disclosed

                                        Conversational AI service, powered initially by LaMDA.

                                        Key Capabilities

                                        • Search integration
                                        • Conversational UI
                                        • Draft generation

                                        Innovations

                                        • Real-time Search Integration
                                        • LaMDA Architecture
                                        • Conversational Search

                                        ERNIE Bot

                                        2023-03-16
                                        Baidu
                                        Architecture not disclosed

                                        Baidu's first ChatGPT competitor, based on ERNIE 3.5.

                                        Key Capabilities

                                        • Conversational AI
                                        • Chinese language
                                        • Multi-turn dialogue

                                        Innovations

                                        • Knowledge Enhanced
                                        • Chinese-first Training
                                        • Baidu Search Integration

                                        Midjourney V5

                                        2023-03-15
                                        Midjourney
                                        Architecture not disclosed

                                        Major leap in photorealism and prompt understanding for AI art generation.

                                        Key Capabilities

                                        • Photorealistic images
                                        • Improved prompts
                                        • Higher resolution

                                        Innovations

                                        • Enhanced Photorealism
                                        • Natural Language Understanding
                                        • Artistic Quality

                                        GPT-4

                                        2023-03-14
                                        OpenAI
                                        Architecture not disclosed

                                        Multimodal model exhibiting human-level performance on various professional and academic benchmarks.

                                        GPQA35.7%
                                        HLE4.2%
                                        SWE-bench1.7%
                                        MMLU86.4%

                                        Key Capabilities

                                        • Multimodal (Image inputs)
                                        • Advanced reasoning
                                        • 32k context window

                                        Innovations

                                        • Native Multimodal Input
                                        • System Card Safety
                                        • Advanced Vision Understanding

                                        Claude 1

                                        2023-03-14
                                        Anthropic
                                        Architecture not disclosed

                                        First version of Claude, focused on helpfulness and safety.

                                        Key Capabilities

                                        • Constitutional AI
                                        • Safety focus
                                        • Helpful assistant

                                        Innovations

                                        • Constitutional AI
                                        • Safety-first Design
                                        • Principle-based Training
                                        Feb

                                        LLaMA

                                        2023-02-24
                                        Meta
                                        Architecture not disclosed

                                        First open-weights large language model from Meta, released to researchers.

                                        Key Capabilities

                                        • Open weights
                                        • 7B to 65B sizes
                                        • Research access

                                        Innovations

                                        • Open-weight LLM
                                        • Efficient Training
                                        • Research Democratization
                                        2022
                                        Nov

                                        ChatGPT

                                        2022-11-30
                                        OpenAI
                                        Architecture not disclosed

                                        Chat interface for GPT-3.5 that took the world by storm.

                                        GPQA28%
                                        HLE2%
                                        SWE-bench3.2%
                                        MMLU70%

                                        Key Capabilities

                                        • Conversational interface
                                        • RLHF tuning
                                        • Broad general knowledge

                                        Innovations

                                        • RLHF Alignment
                                        • Conversational Interface
                                        • Human Preference Learning
                                        Aug
                                        Stability AI
                                        Architecture not disclosed

                                        Revolutionary open-source text-to-image model that democratized AI art.

                                        Key Capabilities

                                        • Text-to-Image
                                        • Open source
                                        • Local inference
                                        • Community ecosystem

                                        Innovations

                                        • Latent Diffusion Model
                                        • Open Source Democratization
                                        • Diffusers Ecosystem
                                        Apr

                                        DALL-E 2

                                        2022-04-06
                                        OpenAI
                                        Architecture not disclosed

                                        Higher resolution image generation with inpainting and variations.

                                        Key Capabilities

                                        • Text-to-Image
                                        • Inpainting
                                        • Variations
                                        • 1024x1024

                                        Innovations

                                        • CLIP Guidance
                                        • Diffusion Model
                                        • Inpainting
                                        2021
                                        Jan

                                        DALL-E

                                        2021-01-05
                                        OpenAI
                                        Architecture not disclosed

                                        AI system that can create realistic images and art from a description in natural language.

                                        Key Capabilities

                                        • Text-to-Image
                                        • Creative variations
                                        • Visual reasoning

                                        Innovations

                                        • Text-to-Image Diffusion
                                        • Visual Reasoning
                                        • Zero-shot Image Generation
                                        2020
                                        Jun

                                        GPT-3

                                        2020-06-11
                                        OpenAI
                                        175B

                                        175 billion parameter language model.

                                        Key Capabilities

                                        • 175B Parameters
                                        • Few-shot learning
                                        • Code generation

                                        Innovations

                                        • Few-shot In-context Learning
                                        • Scale Emergence (175B)
                                        • Near-human Text Generation
                                        2019
                                        Feb

                                        GPT-2

                                        2019-02-14
                                        OpenAI
                                        1.5B

                                        Language model that was initially deemed too dangerous to release.

                                        Key Capabilities

                                        • 1.5B Parameters
                                        • Coherent text generation
                                        • Zero-shot learning

                                        Innovations

                                        • Scaled Transformer Architecture
                                        • Zero-shot Task Transfer
                                        • WebText Dataset