Macro Overview
How are thrones calculated?
Up to 2025, the 👑 LLM throne went to the highest mean of four benchmarks (GPQA, HLE, SWE-Bench Verified, MMLU). From 2026 on a model enters the race with a coding signal — SWE-Bench Pro or Terminal-Bench 2.x — plus GPQA and at least three of the four current-era values (GPQA, SWE-Bench Pro, Terminal-Bench 2.x, MMLU Pro). Scores are min-max normalised across the race pool before they are averaged, so a model gains nothing by omitting the benchmark it would score worst on.
A challenger has to clear the incumbent by 1 normalised point to take the crown, and only a model released after the incumbent can challenge. Inside that band the crown stays put and the line shows a dashed stub to a hollow dot — a live near-tie, not a handover. The 🏆 Coding throne is the same construction over the two coding signals alone (Verified pre-2026).
The 🎬 Video and 🎨 Image thrones use community Elo from the LMArena Text-to-Video and Text-to-Image arenas — the highest-rated model we track wins (with a 95% CI tie-break), since no harmonised academic benchmark exists across video/image generators. Models that aren't on the arena (e.g. Midjourney) can't hold the throne.
LLM and Coding thrones are derived from our own registry — every benchmark score is entered and sourced by hand, with no automated sync. Image and Video thrones are synced from the LMArena leaderboard dataset. Full write-up on the methodology page; for cross-vendor comparisons see artificialanalysis.ai.
Latest model tracked: Sep 10, 2026
DeepSeek-V4.1-Flash
2026-09-10Natively multimodal successor to DeepSeek-V4-Flash-0731 and the first model of DeepSeek's Causal Encoder-Decoder family: 40 layers split into a 20-layer causal encoder and a 20-layer decoder, a 552B backbone plus 196B of sparsely-accessed Engram memory, activating 8B parameters per token during prefill and 16B during decode. 1M-token context, MIT license, open weights. FP4 main KV caching at 890 bytes/token and DSpark speculative decoding drive a reported 409.5 tokens/s end-to-end, and DeepSeek prices it at roughly a quarter of V4-Pro.
Key Capabilities
- 552B backbone MoE, 8B active (prefill) / 16B active (decode)
- Native multimodal image and text input
- 1M token context window
- 384 routed experts (1 shared, 6 routed per token)
- MIT license, open weights
Innovations
- ★Causal Encoder-Decoder (CED) architecture
- ★Compressed Sparse Attention 2 with hierarchical sparse indexer
- ★196B sparsely-accessed Engram memory
- ★FP4 main KV cache at 890 bytes/token
- ★DSpark speculative decoding
- ★Single-Pass mHC
GPT-6 Astra
2026-09-03OpenAI's GPT-6 generation flagship, launched 2026-09-03 in a phased rollout: participants in OpenAI's cybersecurity access program first, then ChatGPT Plus, Pro, Business and Enterprise, the API (model id gpt-6-astra), Azure and AWS Bedrock. 1.05M-token context window, 128K max output, knowledge cutoff 2026-04-30, reasoning effort selectable from none to max. First OpenAI model classified 'Critical' for cybersecurity under the Preparedness Framework: the general release ships with offensive-cyber behaviour constrained, and a fuller cyber variant is limited to vetted defenders. Priced at $10/M input and $50/M output ($1.00 cached input, $12.50 cache write); requests above 272K input tokens are billed at 2x input and 1.5x output for the whole request, and a Fast mode costs 2x. OpenAI's launch table reports Terminal-Bench 4.0, Terminal-Bench-Science, OSWorld 2.0, FrontierMath Tier 4, DeepSWE v1.1 and HLE with tools, but no SWE-bench Pro, Terminal-Bench 2.x, MMLU-Pro or no-tools HLE figure, so the tracked fields here come from independent runs (Artificial Analysis, Vals AI, ARC Prize) where one exists.
Key Capabilities
- Frontier reasoning
- Agentic coding
- Computer use
- Browser use
- Tool use / function calling
- Cybersecurity tasks (vulnerability discovery, exploit development)
- Vision input
Innovations
- ★API: gpt-6-astra
- ★1.05M-token context window
- ★Preparedness Framework 'Critical' cyber classification with a gated cyber variant
- ★Reasoning effort none to max, plus Fast mode at 2x price
- ★Cache reads at 10% of input price
Gemini 3.8 Flash
2026-09-02Third Flash-tier release in six weeks and, per the model card, based on Gemini 3.7 Flash rather than a new training run. Google calls it its most intelligent Flash model yet, with the gains concentrated in software engineering and agentic knowledge work, and describes the model as spending extra reasoning steps and iterative tool calls on hard tasks instead of answering early. 1M-token context, 64K output, knowledge cutoff March 2026, three effort levels. Live at launch in the Gemini app, AI Mode, AI Studio, the Gemini API, Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported (model card, September 2026): DeepSWE v1.1 73.7, Terminal-bench 2.1 89.4, Terminal-bench 4.0 19.1 (a newer generation, tracked on its own axis), HLE-Verified 54.9, Vals Finance Agent v2 61.4, Harvey's Legal Agent Benchmark 10.0, OSWorld-2.0 59.0, CharXiv Reasoning 86.2, LABBench2 86.2, GDPval-AA v2 Elo 1545; Google's API documentation adds SWE-Bench Pro 61.6 and SWE-Atlas 51.9. Introductory API pricing of $0.75 / $3.75 per MTok until 2026-12-31, then $1.50 / $7.50. Announced together with Gemini 3.8 Flash Cyber.
Key Capabilities
- Multimodal input (text, image, audio, video)
- Agentic workflows and coding
- 1M-token context, 64K output
- Configurable effort levels (low / medium / high)
- Available via Gemini API, AI Studio, Antigravity and the Gemini Enterprise Agent Platform
Innovations
- ★Iterates on Gemini 3.7 Flash three weeks after its release (model card: 'based on Gemini 3.7 Flash')
- ★Extra reasoning steps and iterative tool calls on complex tasks, at the cost of more output tokens at higher effort levels
Gemini 3.8 Flash Cyber
2026-09-02Cybersecurity-specialised variant of Gemini 3.8 Flash and successor to Gemini 3.5 Flash Cyber, built to discover, validate and patch software vulnerabilities autonomously across codebases in 20 programming languages. Not a public model: access runs through Google's new Fairwind Program, which pairs the model with the CodeMender harness for government agencies and national cyber authorities, critical-infrastructure operators (healthcare, telecom, energy, finance), maintainers of widely used software platforms and vetted security partners, under operational requirements such as multi-factor authentication and access limited to internal security staff. Google says it will not be generally released. Vendor-reported, none of it a tracked academic benchmark: CyberGym (autonomous vulnerability discovery) above 3.5 Flash Cyber and larger frontier models, no figure published; a success rate above 70% on an internal 20-language vulnerability-discovery benchmark; CWE-Bench (Collinear) patching at 47.2% pass@1; and 2.6x more correct Chrome patches than the best commercial models in the Chrome Security team's own evaluation. No standard academic benchmarks were reported.
Key Capabilities
- Vulnerability discovery, validation and patching across 20 programming languages
- Runs inside the CodeMender harness
- Restricted access: Fairwind Program only (governments, critical infrastructure, core software maintainers, vetted security partners)
Innovations
- ★Cybersecurity-specialised variant of Gemini 3.8 Flash; successor to Gemini 3.5 Flash Cyber
- ★First model distributed through Google's Fairwind Program, a gated access scheme with operational requirements instead of a general release
Muse Spark 1.3
2026-09-02Fourth Muse Spark release in five months, live at launch in Muse Code and the Meta Model API. Meta frames the update as an efficiency change as much as a capability one: for the same agentic work it reports about 20% fewer tool calls and about 25% fewer tokens, with fewer turns and less verbose output. 1M-token context, multimodal input; xhigh reasoning effort is the generally available tier, a max tier is in limited preview for Meta partners. Vendor-reported (Meta launch scorecard, September 2026): Terminal-Bench 2.1 88.8, DeepSWE v1.1 75.4, SWEAtlas CodeBase QnA 59.4, MRCR 256K-512K 98.5 and MRCR 512K-1M 98.1. API pricing unchanged from 1.1 and 1.2 at $1.25 / $4.25 per million input/output tokens with cached input at $0.15; the contributor tier (over 90% discount in exchange for training on prompts and completions) continues.
Key Capabilities
- Agentic coding
- Long-horizon tool use
- Multimodal input
- 1M-token context
- Reasoning effort control (xhigh; max in limited preview)
Innovations
- ★Token and tool-call efficiency as a stated release goal, at unchanged price and context
- ★Fourth frontier release from Meta Superintelligence Labs in five months
Claude Fable 5.1
2026-09-01Successor to Claude Fable 5 in the Mythos-class tier, at the same $10/M input and $50/M output pricing but with cache reads cut to $0.25/M (2.5% of the input price instead of 10%). Launched alongside Claude Mythos 5.1, the same underlying model with safeguards lifted for Project Glasswing participants (tracked separately). Safety classifiers remain in place; refused requests can fall back server-side to Claude Opus 4.8 or Claude Opus 5. Adaptive thinking is always on, forced tool use is no longer supported, and all text output carries Anthropic's statistical watermark. Vendor-reported benchmarks: Terminal-Bench 4.0 55.8 (Fable 5: 42.0, Opus 5: 52.3), Terminal-Bench-Science 0.1 52.6, HLE 60.9 without tools (65.0 with tools), OSWorld 2.0 41.7 strict, CursorBench 3.2.0 73.4, AutomationBench 31.4.
Key Capabilities
- Autonomous long-horizon tasks
- Agentic coding
- Multistep research
- Computer use
- Vision input
- Tool use
Innovations
- ★API: claude-fable-5-1
- ★Cache reads at 2.5% of input price
- ★Content provenance (text watermark, C2PA for files)
- ★Per-message effort control
Claude Mythos 5.1
2026-09-01Same underlying model as Claude Fable 5.1, offered by invitation only to Project Glasswing participants and the Cyber and Life Sciences verification programs, with the safety classifiers and fallback routing that govern Fable 5.1 lifted. Successor to Claude Mythos 5. Shares Fable 5.1's specifications and pricing ($10/M input, $50/M output, 1M context). API ID claude-mythos-5-1 on the Claude API, Bedrock, Google Cloud and Foundry; no public access. Anthropic reports Terminal-Bench 4.0 60.9 against 55.8 for Fable 5.1 — the gap is the cost of the safeguard interventions — and HLE 60.9 without tools (65.0 with tools) for both models. Independent evaluators measure Fable 5.1, not this variant, so no independent values are carried over.
Key Capabilities
- Autonomous long-horizon tasks
- Agentic coding
- Cyber defense
- Multistep research
- Computer use
- Vision input
Innovations
- ★API: claude-mythos-5-1
- ★Project Glasswing restricted access
- ★Safeguards lifted for verified defenders
Atlas
2026-09-01World Labs' omni world model, pretrained from scratch to operate natively on text, images, video and 3D as a multimodal autoregressive diffusion transformer with every input in one shared spatial context. The headline capability is camera-controlled video, up to one minute at 1440p from one or more reference images, with the camera path supplied as a geometric input rather than described in a prompt. The same model reconstructs scenes into point clouds and 3D Gaussian splats, works with depth maps and camera poses, reframes multi-camera footage and supports parts of a real-to-sim robotics workflow. Early access for selected partners via a request form; no pricing, API or public date announced. World Labs reports that third-party human raters preferred Atlas over rival video models in 75 to 94 percent of pairwise trials depending on the competitor, a vendor-commissioned preference study recorded here as a note, not as a score.
Key Capabilities
- Camera-controlled video up to one minute at 1440p
- Text, image, video and 3D inputs in one spatial context
- 3D reconstruction to point clouds and Gaussian splats
- Real-to-sim workflows for robotics
Innovations
- ★Single omni model for generation, reconstruction and simulation
- ★Camera path as geometric input instead of prompt text
Hy4-preview
2026-08-28Preview of Tencent's next-generation Hunyuan flagship: a 770B-parameter MoE activating 49B per token, with a 1M-token context window. Open-weight under Apache 2.0 with an FP8 checkpoint alongside BF16, tuned for coding agents, complex tool-use workflows and productivity tasks, and trained on domain data from Tencent's software, gaming and finance teams.
Key Capabilities
- 770B total / 49B active MoE
- 1M token context window
- Coding-agent and tool-use focus
- Apache 2.0 license
Innovations
- ★Tencent-internal software, gaming and finance domain data
- ★FP8 checkpoint shipped alongside BF16
Gemini Omni 1.1 Flash
2026-08-27Successor to Gemini Omni Flash (2026-06-30) for video generation and editing, released as the model id gemini-omni-1.1-flash. Takes any combination of text, image, audio and video as input and returns video. Scene extension now conditions on up to 10 seconds of preceding footage instead of the tail of the clip and continues in 10-second increments up to 40 seconds total; a first-and-last-frame control renders a single continuous shot between two supplied frames, and a video reference of up to 3 seconds carries character and style across generations. Native output resolution is 720p, with 1080p and 4K delivered as upscales, plus a new 360p draft mode that Google reports as up to 60% faster at a third of the 720p cost. Per-second API pricing: $0.03 (360p), $0.10 (720p), $0.15 (1080p), $0.30 (4K). Live via the Gemini API, Google AI Studio, Flow and the Gemini Enterprise Agent Platform; scene extension is rolling out to Google AI Plus, Pro and Ultra subscribers in the Gemini app globally. The predecessor's gemini-omni-flash-preview endpoint is scheduled for deprecation on 2026-09-30.
Key Capabilities
- Text-to-video
- Image/audio/video-to-video (any-to-any input)
- Conversational multi-turn editing
- Scene extension up to 40s in 10s increments
- First-and-last-frame shot control
- Video reference up to 3s for character/style consistency
- 720p native output, 1080p and 4K upscale
- 360p draft mode
Innovations
- ★Scene extension conditioned on up to 10s of prior footage, extending to 40s total
- ★First-and-last-frame control for a single continuous shot
- ★360p draft mode at a third of the 720p per-second price
- ★4K upscaled export at $0.30 per second
DeepSeek-V4-Flash-Vision-Exp
2026-08-21Experimental vision variant of DeepSeek-V4-Flash-0731, live on the DeepSeek API platform as model 'deepseek-v4-flash-vision-exp'. Same 284B total / 13B active MoE and 1M-token context as the text build, extended to image input (text output only); images are billed as up to 384 tokens each and can be passed as base64 through Chat Completions, Messages and the Responses API. DeepSeek describes text capability — agents, reasoning, world knowledge — as unchanged versus V4-Flash and reports the gains on agent benchmarks that require visual understanding: Terminal-Bench 2.1 83.9 and DeepSWE 59.3. Unlike the rest of the V4 line this is a closed API release — no open weights, no Hugging Face repository at launch, and the 'Exp' tag marks it as experimental.
Key Capabilities
- 284B total / 13B active MoE
- Image input, text output
- 1M token context window
- Multimodal agent tasks
- Chat Completions, Messages and Responses API
- Images billed at up to 384 tokens each
Innovations
- ★First vision-enabled model of the DeepSeek V4 line
- ★Vision added on top of the V4-Flash-0731 checkpoint at unchanged parameter count
- ★API-only release — first V4 build without open weights
GLM-5.3
2026-08-14Post-training-only update to GLM-5.2: Z.ai leaves the 753B MoE base model unchanged and attributes the capability gains entirely to scaled post-training. Weights landed on Hugging Face on 2026-08-28, fourteen days after launch, once Z.ai's safety review completed — but under a bespoke GLM-5.3 License instead of the MIT license GLM-5.2 shipped under: use, modification, distribution, sublicensing, sale, deployment and fine-tuning are all permitted, but a company with more than $10B aggregate revenue over any 12 consecutive months must pass a Z.ai security review before hosting the model commercially. Individual users and smaller companies are unaffected. Z.ai reports its results on a new generation of agentic benchmarks rather than the axes tracked here (maximum thinking effort): Terminal-Bench 3.0 28.3 (GLM-5.2: 4.6), DeepSWE v1.1 66.9 (46.2), Agents' Last Exam 28.5 (23.8), CyberGym 84.5, plus AutomationBench, GDPVal-AA v2 and HLE with tools. Only Terminal-Bench 3.0 has a column here (its own axis, separate from the 2.x scale); the rest do not map onto the tracked benchmarks, and Z.ai published no GPQA, MMLU-Pro or SWE-bench Pro table.
Key Capabilities
- 753B total / ~40B active MoE (base unchanged from GLM-5.2)
- 1M-token context
- Agentic / long-horizon coding
- Cybersecurity workflows
- Open weights under the custom GLM-5.3 License (revenue-tiered)
Innovations
- ★Capability gain from post-training scaling alone, base model unchanged
- ★Revenue-gated open-weight license replacing MIT (security review above $10B revenue)
DeepSeek-V4-Pro-0813
2026-08-13General-availability build of DeepSeek-V4-Pro, replacing the April preview. Same 1.6T total / 49B active MoE architecture as the preview — re-post-trained, with the gains concentrated in agentic and software-engineering tasks. The deepseek-v4-pro API endpoint now points to this build. 1M-token context, MIT license.
Key Capabilities
- 1.6T total / 49B active MoE
- 1M token context window
- Three reasoning effort modes (incl. Think Max)
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★Re-post-training of the April preview checkpoint
Gemini 3.7 Flash
2026-08-13Workhorse successor to Gemini 3.6 Flash, shipped three weeks after it. Google states the model was not trained from scratch — it replaces the predecessor via algorithmic improvements and user feedback. Same 1,048,576-token context window as 3.6 Flash, multimodal input (text, image, audio, video, files). Live through the Gemini API in AI Studio, Android Studio, Google Antigravity and the Gemini Enterprise Agent Platform. Vendor-reported coding deltas against 3.6 Flash: DeepSWE v1.1 49.0 to 65.3, FrontierCode 1.1 Main 34.4 to 43.6. Introductory API pricing of $0.75 / $3.75 per MTok runs until 2026-12-31, after which Google lists $1.50 / $7.50.
Key Capabilities
- Multimodal input (text, image, audio, video, files)
- Agentic workflows and coding
- 1M-token context
- Available via Gemini API, Android Studio and Antigravity
Innovations
- ★Successor built by replacing 3.6 Flash through algorithmic improvements rather than a from-scratch training run
Qwen3.8-2.4T-A95B
2026-08-12Open-weight checkpoint of the Max-class flagship and the first Qwen-Max-class model with downloadable weights: 2.4T total / 95B active sparse MoE, published on Hugging Face and ModelScope alongside an FP8 build and served via API. It is not the hosted Qwen3.8-Max — the checkpoint is text-only (no vision or video input) and always reasons, thinking mode cannot be disabled. Native context is 262,144 tokens, extensible to roughly 1M via YaRN. Shipped under the bespoke Qwen3.8-Max License rather than Apache 2.0. In FP8 the weights occupy about 2,325 GiB, so self-hosting takes a 16-GPU class deployment.
Key Capabilities
- 2.4T total / 95B active parameters (sparse MoE)
- 262K native context, ~1M via YaRN
- Text-only input and output
- Always-on thinking (cannot be disabled)
- 128K max output tokens
- FP8 and BF16 checkpoints, vLLM / SGLang support
Innovations
- ★First Qwen-Max-class model released with open weights
- ★2.4T-parameter open-weight MoE
- ★Fine-grained FP8 block quantization (block size 128)
Grok 4.6
2026-08-12Reasoning-first flagship succeeding Grok 4.5, with a 500K-token context window and text + image input. Same $2 / $6 per million input/output tokens as Grok 4.5, with a 75% cache-hit discount.
Key Capabilities
- 500K token context window
- Agentic coding
- Reasoning-first
- Text and image input
Innovations
Muse Glimmer
2026-08-1030B dense agentic model distilled from Muse Spark 1.2 and released under Apache 2.0 — Meta's first open-weight release since the Llama 4 family. Served through the API as well as downloadable: at 4-bit the checkpoint stays under 20 GB, so the full setup including KV cache and perception encoder fits a 24–32 GB envelope on a single consumer GPU or Mac. A separate perception encoder handles image input; block-level speculative decoding keeps latency inside a real agent loop. 131K-token context.
Key Capabilities
- 30B dense parameters
- Multimodal input via separate perception encoder
- Agentic tool calling with planning and failure recovery
- Controllable reasoning effort levels
- 131K-token context
- Apache 2.0
Innovations
- ★Distilled from Muse Spark 1.2
- ★Block-level speculative decoding for agent-loop latency
- ★4-bit K-Quant checkpoint under 20 GB
- ★Meta's first open-weight release since Llama 4
Muse Spark 1.2
2026-08-05Coding-focused update of the Muse Spark frontier family, released alongside and co-trained with the Muse Code terminal agent. Third Meta Superintelligence Labs model in four months. Meta evaluates it inside its own Muse Code harness at xhigh reasoning effort. Same API pricing as Muse Spark 1.1 at $1.25 / $4.25 per million input/output tokens, with cached input at $0.15; a contributor tier cuts the rate by over 90% in exchange for permission to train on prompts and completions. Rate limits reach 3.000 requests and 4M tokens per minute per team.
Key Capabilities
- Agentic coding
- Long-horizon tool use
- Multimodal input
- Multi-agent orchestration
- Reasoning effort control (xhigh)
Innovations
- ★Co-trained with the Muse Code agent harness
- ★Contributor tier (data-sharing discount)
Muse Code
2026-08-05Meta’s first coding agent: a terminal-based harness for macOS and Linux, released in beta and powered by Muse Spark 1.2. Runs a simple agent loop plus persistent async background agents that stay alive for a whole session and accumulate context instead of restarting per task; larger tasks are split onto parallel subagents in isolated git worktrees so the working copy stays untouched. Meta reports one run optimising Nvidia Hopper GPU kernels across more than 1.000 tool calls in sessions of up to 24 hours. Pay-as-you-go at Muse Spark 1.1 API rates, or a contributor tier at over 90% discount in exchange for training on prompts and completions.
Key Capabilities
- Terminal coding agent
- Repository-scale planning, editing and validation
- Persistent async background agents
- Parallel subagents in isolated worktrees
- Long-running sessions (24h+)
Innovations
- ★Session-persistent background agents
- ★Contributor tier (data-sharing discount)
Qwen3.8-Max
2026-08-02General-availability build of the flagship Max model, two weeks after the WAIC preview. 2.4T-parameter sparse Mixture-of-Experts, text/image/video in and text out, 1M-token context (max ~991K input, ~983K with thinking enabled; 65,536 output tokens). Served via Alibaba Cloud Model Studio with an OpenAI- and DashScope-compatible API at $2 / $6 per MTok flat across the full context. The active-parameter count is still undisclosed and no benchmark table has been published. Alibaba announced open weights for Qwen3.8-Max plus a smaller Qwen3.8-27B checkpoint for the following week — the first Qwen-Max-class model to go open-weight. Both have since shipped and are tracked as separate local entries: Qwen3.8-2.4T-A95B (2026-08-12, text-only and always-thinking, custom Qwen3.8-Max License) and Qwen3.8-27B (2026-08-14, Apache 2.0). The open checkpoint is not identical to this hosted API — it drops the vision path — so the benchmark values here, measured against the API, are not carried over to it.
Key Capabilities
- 2.4T total parameters (sparse MoE)
- 1M token context window
- Multimodal input (text, image, video), text output
- 65,536 max output tokens
- OpenAI- and DashScope-compatible API
- Open weights released as separate checkpoints (2.4T-A95B, 27B)
Innovations
- ★2.4T Sparse MoE
- ★First open-weight release of a Qwen-Max-class model
- ★Flat pricing across the full 1M-token context
DeepSeek-V4-Flash-0731
2026-07-31General-availability build of DeepSeek-V4-Flash, replacing the April preview. Same 284B total / 13B active MoE architecture and size as the preview — only re-post-trained, with the gains concentrated in agentic tool use. 1M-token context, MIT license. On 2026-08-21 DeepSeek added a vision variant on top of this checkpoint, DeepSeek-V4-Flash-Vision-Exp — API-only, without open weights. Superseded on 2026-09-10 by DeepSeek-V4.1-Flash, which moves to a Causal Encoder-Decoder architecture rather than re-post-training this checkpoint.
Key Capabilities
- 284B total / 13B active MoE
- 1M token context window
- Three reasoning effort modes (incl. Think Max)
- Responses API support
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★Agent-focused re-post-training on the preview checkpoint
Gemini Robotics 2
2026-07-30Second generation of Google DeepMind's robotics stack, extending control from the upper body to the whole body: a humanoid can walk, crouch, lean and manipulate in one coordinated motion, and several robots can collaborate on a task. Ships as three models — Gemini Robotics-ER 2 (reasoning, available in Google AI Studio and in private preview on the Gemini Enterprise Agent Platform), plus the VLA and an on-device variant for early-access partners.
Key Capabilities
- Whole-body humanoid control
- Advanced dexterity
- Multi-robot collaboration
- On-device variant
Innovations
- ★Whole-body control including legs, not just manipulation
- ★Multi-robot task coordination
Claude Opus 5
2026-07-24Opus-tier Claude flagship for agentic coding, long-horizon tool use and computer use, released as the new top model in the Claude apps and the API. Priced at $5 / $25 per million input/output tokens.
Key Capabilities
- Agentic coding
- Computer use
- Extended thinking
- Tool use
- Vision input
Innovations
- ★API: claude-opus-5
- ★Effort control
FLUX 3
2026-07-23Unified multimodal flow model that generates image, video and audio from a single backbone and can be extended to predict robot actions. Rollout is staged: FLUX 3 Video with native synchronised audio (clips up to 20 seconds, aspect ratios from 9:16 to 21:9) and the action model are in application-based early access, image generation is announced for the following weeks, and an open-weight FLUX 3 Dev backbone for later in the year. No public API tier, pricing, parameter count or benchmark methodology published at launch. FLUX 3 Video now ranks #2 in the Text-to-Video Arena at ~1,496 Elo, behind Gemini Omni Flash (~1,512) and ahead of Dreamina Seedance 2.0 (~1,478) — the first independent quality signal for the model.
Key Capabilities
- Text-to-video (up to 20s)
- Native synchronised audio
- Image-to-video
- Video-to-video from reference clip
- Keyframe-to-video transitions
- Robot action prediction
Innovations
- ★Single flow backbone for image, video, audio and action
- ★Video and native audio generated jointly
- ★Text-to-Video Arena #2 at intake (~1,496 Elo)
Qwen-Image-3.0
2026-07-21Third-generation text-to-image model from Alibaba's Qwen team, focused on dense text and layout rendering. Accepts prompts up to 4,500 tokens, renders in-image text in 12 languages and 20+ fonts down to ~10px, and produces complex multi-element layouts (newspapers, storyboards, infographics, UI mockups, knowledge graphs) in a single pass, generating up to 9 images at once. Unlike the earlier Apache 2.0 releases (Qwen-Image 1.0, Aug 2025, and 2.0, each with a technical report), 3.0 shipped without open weights, benchmarks or a technical report. Available via chat.qwen.ai, Qwen Studio and Alibaba's API platform; API pricing not disclosed at launch.
Key Capabilities
- Text-to-Image
- Prompts up to 4,500 tokens
- In-image text rendering (12 languages, 20+ fonts, ~10px)
- Single-pass complex layouts
- Up to 9 images per generation
Innovations
- ★Long-prompt layout generation (4.5K tokens)
- ★High-density multilingual text rendering
- ★Closed release (no open weights, unlike prior Qwen-Image versions)
Gemini 3.6 Flash
2026-07-21Workhorse model in the Gemini 3 series, based on Gemini 3.5 Flash, delivering better coding, knowledge work and multimodal performance at improved token-efficiency (Google reports ~17% fewer output tokens than 3.5 Flash). 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch across the Gemini app, AI Studio and Gemini API. Vendor-reported (model card): SWE-Bench Pro 58.7, Terminal-Bench 2.1 78.0, OSWorld-Verified 83.0, MLE-Bench 63.9, CharXiv Reasoning (with tools) 89.4, GDM-MRCR v2 (128k) 91.8. Google's launch benchmark table also lists GPQA Diamond 90.4 (unchanged vs 3.5 Flash). Announced alongside 3.5 Flash-Lite and 3.5 Flash Cyber.
Key Capabilities
- Multimodal input (text, image, audio, video)
- Agentic workflows and coding
- 1M-token context
- GA at launch (2026-07-21)
Innovations
- ★~17% fewer output tokens than Gemini 3.5 Flash at comparable or better quality (token-efficiency)
Gemini 3.5 Flash-Lite
2026-07-21Cost-efficient, low-latency tier of the Gemini 3 series, based on Gemini 3.1 Flash-Lite and optimized for high-volume, latency-sensitive tasks (translation, classification, document processing) as well as agentic workflows. 1M-token context, 64K output, knowledge cutoff March 2026. GA at launch and rolling out to Google Search. Vendor-reported (model card): SWE-Bench Pro 54.2, Terminal-Bench 2.1 54.0, OSWorld-Verified 74.0, MLE-Bench 39.2, CharXiv Reasoning (with tools) 76.5, GDM-MRCR v2 (128k) 72.2. Reported output speed ~350 tokens/s.
Key Capabilities
- Multimodal input (text, image, audio, video)
- High-throughput, low-latency tasks
- Agentic workflows
- 1M-token context
- Rolling out to Google Search
Innovations
- ★Output speed ~350 tok/s at $0.30 / $2.50 per 1M tokens (cost-efficient high-volume tier)
Gemini 3.5 Flash Cyber
2026-07-21Security-specialized model fine-tuned from Gemini 3.5 Flash to identify, validate and patch software vulnerabilities at low cost per token. Deployed inside Google's CodeMender security agent, which calls it many times in rapid succession to explore alternative execution paths across a codebase. Availability is limited: an initial pilot for governments and trusted partners via CodeMender, not general API access. Google reports internal results of 55 confirmed unique vulnerabilities on Chrome's V8 JavaScript engine (vs 47 for standard 3.5 Flash and 36 for Claude Opus 4.6), a 42% improvement on long-range multi-turn cyber benchmarks over the Flash 3 predecessor, and frontier-competitive performance on CyberGym within CodeMender. No standard academic benchmarks were reported.
Key Capabilities
- Vulnerability detection, validation and patching
- Runs inside the CodeMender security agent
- Limited pilot: governments and trusted partners only
Innovations
- ★Cybersecurity-specialized fine-tune of Gemini 3.5 Flash for automated vulnerability discovery and patching
Qwen3.8-Max-Preview
2026-07-19Preview of Alibaba's flagship Max model, unveiled at WAIC Shanghai. 2.4T-parameter sparse Mixture-of-Experts, natively multimodal (text, images, video, documents) with a 1M-token context window (inherited from Qwen3.7-Max). Active-parameter count, full benchmark suite and license were not disclosed at preview. Runs at 10% of standard pricing during the preview via Token Plan / Qoder; an open-weight release was announced as forthcoming.
Key Capabilities
- 2.4T total parameters (sparse MoE)
- 1M context window
- Natively multimodal (text, image, video, docs)
- Preview via Token Plan / Qoder
- Open weights planned
Innovations
- ★2.4T Sparse MoE
- ★Multimodal Text/Image/Video/Docs
- ★Open-Weight Max Release Planned
Kimi K3
2026-07-162.8T MoE (~50B active, 16 of 896 experts) native multimodal flagship with a 1M-token context window. Launched via API, app and playground on 2026-07-16; open weights scheduled for 2026-07-27.
Key Capabilities
- 2.8T total / ~50B active MoE (896 experts, 16 routed)
- 1M token context window
- Natively multimodal (text, image, video)
- Vision-in-the-loop (screenshot inspect + code edit)
- Open weights scheduled 2026-07-27
Innovations
- ★Kimi Delta Attention
- ★2.8T Open-Weight MoE (largest open-weight at launch)
- ★Vision-in-the-Loop Agent Feedback
- ★1M-Token Native Context
Muse Spark 1.1
2026-07-09Second Meta Superintelligence Labs model and the launch vehicle for the paid Meta Model API. Multimodal reasoning model built for agentic work: takes text, image, video, audio and PDF input, returns text, with a 1M-token context window. Runs as main agent that plans and delegates or as a subagent, and generalises zero-shot to new tools, MCP servers and custom skills. Free for consumers in the Meta AI app (Thinking mode); developer preview US-only at launch. $1.25 / $4.25 per million input/output tokens.
Key Capabilities
- Multimodal input (text, image, video, audio, PDF)
- 1M token context window
- Agentic tool use
- Multi-agent orchestration (main agent / subagent)
- MCP server support
- Reasoning
Innovations
- ★Meta Model API (first paid Meta inference API)
- ★Zero-shot generalisation to unseen tools, MCP servers and skills
Grok 4.5
2026-07-08Coding- and agentic-focused flagship from xAI, trained alongside Cursor. Available in Grok Build, Cursor and the console at $2 / $6 per million input/output tokens.
Key Capabilities
- Agentic coding
- Reasoning-first
Innovations
Reve 2.1
2026-07-08Iteration on Reve's layout-first 4K text-to-image model, with improved prompt understanding, world knowledge and in-image text rendering (including foreign scripts). Retains the layout-as-prompt architecture and native 4K output.
Key Capabilities
- Layout-first generation
- Native 4K (4096×4096) output
- Addressable / editable elements
- Text rendering in images
- Foreign-script text rendering
Innovations
- ★Layout-as-prompt (position + size + local description)
- ★Native 4096×4096 output
Seedream 5.0 Pro
2026-07-08Pro tier of ByteDance's Seedream 5.0 image family (above Seedream 5.0 and 5.0 Lite). Adds reasoning- and online-search-guided generation, separable-layer editing, and text rendering in 10+ languages, aimed at high-density infographics and professional editing.
Key Capabilities
- Text-to-Image
- Reasoning-guided generation
- Online-search-guided generation
- Separable-layer editing
- Text rendering in 10+ languages
Innovations
- ★Reasoning + online search for image generation
- ★Layer-separable editing
- ★Multilingual text rendering (10+ languages)
Muse Image
2026-07-07Meta Superintelligence Labs' first image-generation model (codename Mango). Operates agentically, invoking search and coding tools, self-refining its outputs and scaling test-time compute, for instruction following, precise editing and multi-reference composition. Available in the Meta AI app, on meta.ai, in Instagram Stories (US) and WhatsApp (limited).
Key Capabilities
- Text-to-Image
- Precise image editing
- Multi-reference composition
- Agentic tool use (search + code)
- Test-time self-refinement
Innovations
- ★First Meta Superintelligence Labs image model
- ★Agentic generation (tool use + self-refinement)
Hy3
2026-07-06Tencent's Hunyuan 3 flagship: a 295B-parameter MoE activating 21B per token (192 routed experts with top-8 routing plus an always-active shared expert), 256K context window and three selectable reasoning-effort modes. Full release under Apache 2.0 — unlike the April preview license, without territorial restrictions.
Key Capabilities
- 295B total / 21B active MoE
- 256K token context window
- Three reasoning-effort modes
- Apache 2.0 license
Innovations
- ★Dense-MoE hybrid with always-active shared expert
- ★Multi-Token-Prediction layers (3.8B)
Claude Sonnet 5
2026-06-30Sonnet-tier Claude model for agentic coding, computer use and tool use, released as the new default in the Claude apps and the API.
Key Capabilities
- Agentic coding
- Computer use
- Extended thinking
- Tool use
- Vision input
Innovations
- ★API: claude-sonnet-5
- ★OSWorld-Verified computer use
- ★Introductory pricing through 2026-08-31
Nano Banana 2 Lite
2026-06-30Faster, lower-cost variant of Nano Banana 2 for rapid text-to-image generation (~4s per image).
Key Capabilities
- Fast text-to-image (~4s)
- Text rendering in images
- Low per-image cost
- Image editing
Innovations
- ★Lite tier of Gemini 3.1 Flash Image
- ★~$0.034 per generated image
- ★Latency-optimized generation
Gemini Omni Flash
2026-06-30Google's cost-efficient multimodal video-generation model with conversational editing; generates and refines video from text, image, audio and video inputs. 720p, up to 10s clips, $0.10/sec via the Gemini API.
Key Capabilities
- Text-to-video
- Image/audio/video-to-video
- Conversational multi-turn editing
- 720p output
- Up to 10s clips
Innovations
- ★Gemini Omni model family
- ★Native conversational video editing
- ★$0.10 per second (Gemini API / AI Studio)
GPT-5.6 Sol
2026-06-26Preview of OpenAI's GPT-5.6 frontier model (Sol tier), part of a Sol/Terra/Luna family in limited preview to trusted partners via Codex and the API. Sol priced at $5/M input, $30/M output tokens. Not yet generally available.
Key Capabilities
- Frontier reasoning
- Agentic coding
- Tool use / function calling
- Computer use
- Cybersecurity tasks
Innovations
- ★Sol / Terra / Luna model family
- ★Limited preview (Codex + API)
- ★Token-efficient agentic workflows
GPT-5.6 Terra
2026-06-26Balanced mid-tier of OpenAI's GPT-5.6 Sol/Terra/Luna family, positioned as the everyday default for interactive and agentic coding. Priced at $2.50/M input, $15/M output tokens. Part of the limited preview from 2026-06-26; generally available 2026-07-09.
Key Capabilities
- Agentic coding
- Tool use / function calling
- Programmatic tool calling (Responses API)
Innovations
- ★Sol / Terra / Luna model family
GPT-5.6 Luna
2026-06-26Lightweight, cost-efficient tier of OpenAI's GPT-5.6 Sol/Terra/Luna family, aimed at smaller and faster tasks. Lowest-cost option in the family at $1/M input, $6/M output tokens. Part of the limited preview from 2026-06-26; generally available 2026-07-09.
Key Capabilities
- Agentic coding
- Tool use / function calling
- Programmatic tool calling (Responses API)
Innovations
- ★Sol / Terra / Luna model family
Doubao 2.1 Pro
2026-06-23Auf der Volcano Engine Force Conference 2026 vorgestelltes Flaggschiff (offiziell Doubao-Seed-2.1 Pro), positioniert am production-level capability threshold. Schwerpunkte: Code-Delivery, Long-Horizon-Agent-Tasks, multimodales Verständnis und Enterprise-Stabilität. Proprietär über Volcano Engine. Unabhängige Benchmark-Werte stehen zum Eintragungszeitpunkt noch aus.
Key Capabilities
- Coding / Engineering Delivery
- Long-Horizon Agent
- Vision Language / Multimodal
- Enterprise-Stabilität
Innovations
- ★Production-Threshold-Flaggschiff
Seedance 2.5
2026-06-23Beta preview of ByteDance's Seedance 2.5 video model, headline feature 30-second one-shot generation. Announced as a global enterprise beta at Volcano Engine 2026; public launch targeted for early July 2026.
Key Capabilities
- Text-to-video
- 30s one-shot generation
- Audio-video joint generation
- Multimodal inputs
Innovations
- ★30-second single-shot video
- ★Enterprise beta (Volcano Engine)
- ★Seedance 2.x architecture
GLM-5.2
2026-06-13753B Mixture-of-Experts (~40B active) open-weight model under MIT license. Introduces the IndexShare attention mechanism (~2.9x lower per-token FLOPs at 1M-token context) and a usable 1M-token context window. Ships with two thinking-effort levels; the highest tier is marketed as 'Max'. Available as open weights on Hugging Face and via the Z.ai API. Vendor-reported benchmarks (maximum thinking effort): GPQA Diamond 91.2, SWE-bench Pro 62.1, Terminal-Bench 2.1 81.0, HLE 40.5 (no tools).
Key Capabilities
- 753B total / ~40B active MoE
- 1M-token context
- Two thinking-effort levels (incl. 'Max')
- Agentic / long-horizon coding
- MIT license
Innovations
- ★IndexShare attention (~2.9x FLOP reduction at 1M-token context)
- ★Usable 1M-token context window
Claude Fable 5
2026-06-09Anthropic's first publicly available Mythos-class model, priced at $10/M input and $50/M output. Launched alongside the gated Claude Mythos 5 (same base model with safeguards lifted for a small set of cyberdefenders and infrastructure providers), which is not publicly available and is not tracked here. A safety system routes a minority of sessions (~5% on average) to Claude Opus 4.8. Vendor-reported benchmarks: SWE-bench Pro 80.3, HLE 59.0 (no tools; 64.5 with tools), Terminal-Bench 2.1 88.0.
Key Capabilities
- Autonomous long-horizon tasks
- Vision input
- Agentic workflows
- Software engineering
Innovations
- ★First publicly available Mythos-class model
- ★Safety routing to Claude Opus 4.8
Reve 2.0
2026-06-03Layout-first 4K text-to-image model from Reve. Represents each element with a position, size and local description for code-like, addressable editing. Native 4096×4096 output.
Key Capabilities
- Layout-first generation
- Native 4K (4096×4096) output
- Addressable / editable elements
- Text rendering in images
Innovations
- ★Layout-as-prompt (position + size + local description)
- ★Native 4096×4096 output
MAI-Code-1-Flash
2026-06-02Microsoft's first in-house coding model, part of the MAI family unveiled at Build 2026 and built on Maia 200 silicon without distillation from other labs. Live in GitHub Copilot and VS Code, and available to developers via OpenRouter, Fireworks and Baseten with tunable weights. Vendor-reported benchmarks: SWE-bench Pro 51.2, GPQA Diamond 84.6 (model card).
Key Capabilities
- Agentic coding
- GitHub Copilot integration
- Low token usage
Innovations
- ★First Microsoft in-house coding model
- ★Trained on Maia 200 silicon
- ★Developer-tunable weights
MAI-Image-2.5
2026-06-02Text-to-image and image-editing model from Microsoft's MAI family, unveiled at Build 2026 and built in-house on Maia 200 silicon. Available through Microsoft Foundry, with a faster Flash variant that is not tracked separately here.
Key Capabilities
- Text-to-image
- Image editing
Innovations
- ★First Microsoft in-house image model
- ★Trained on Maia 200 silicon
MAI-Thinking-1
2026-06-02Microsoft AI's first flagship reasoning model, a medium-sized model in the first-party MAI family launched at Build 2026. Distributed via OpenRouter, Fireworks and Baseten.
Key Capabilities
- Extended reasoning
- Agentic coding
- Mathematical reasoning
- Tool use
Innovations
- ★First-party Microsoft reasoning flagship
- ★MAI model family (Build 2026)
Claude Opus 4.8
2026-05-28Anthropic's flagship upgrade to Opus 4.7, released at the same standard price ($5/M input, $25/M output). 1M-token input context with up to 128K output tokens. Adds a Fast mode running at 2.5x output speed ($10/M input, $50/M output) alongside the standard tier. New 'dynamic workflows' tool for Claude Code, effort control in claude.ai and Cowork, and a Messages API that now accepts system entries mid-conversation. Vendor-reported benchmarks: GPQA Diamond 93.6, SWE-bench Verified 88.6, SWE-bench Pro 69.2, Terminal-Bench 2.1 74.6. Artificial Analysis Intelligence Index v4.0: 61.
Key Capabilities
- 1M-token input context window
- Up to 128K output tokens
- Vision input
- Computer use
- Agentic workflows
Innovations
- ★Dynamic Workflows (Claude Code)
- ★Effort Control (claude.ai & Cowork)
- ★Mid-Conversation System Entries (Messages API)
- ★Fast Mode (2.5x speed tier)
Qwen3.7-Max
2026-05-20Proprietary flagship in Alibaba's Qwen Max line. Closed weights — the Plus variant is open-sourced while Max stays closed, consistent with Alibaba's Max/Plus split since Qwen2.5-Max. 1M-token context (up from 256K in the Qwen3.6-Max preview), text-only input and output. MoE architecture details are not officially disclosed. Previewed on Arena AI on 2026-05-14, formally launched at the 2026 Alibaba Cloud Summit on 2026-05-20. Alibaba-reported benchmarks (verified post-launch via AA Intelligence Index v4.0 reproduction): GPQA Diamond 92.4, HLE 41.4, SWE-bench Verified 80.4, Terminal-Bench 2.0 69.7. AA Intelligence Index v4.0: 57 (initial Index run reported 56.6).
Key Capabilities
- 1M-token context window
- Text-only input and output
- Closed weights
Innovations
Command A+
2026-05-20218B sparse Mixture-of-Experts (~25B active) open-weight flagship under Apache 2.0, runnable on as few as 2x H100 GPUs. Enterprise- and RAG-focused with strong multilingual coverage. Available as open weights and via Cohere's API. Early independent benchmarks (Artificial Analysis): GPQA Diamond ~76, HLE ~11, MMMU-Pro 63, Intelligence Index ~37 — flagged as estimates, AA's full evaluation is still ongoing.
Key Capabilities
- 218B sparse MoE / ~25B active
- Apache 2.0 open weights
- Multilingual
- Enterprise / RAG focus
- Runs on 2x H100
Innovations
- ★Apache 2.0 open-weight enterprise flagship
Gemini 3.5 Flash
2026-05-19Fast-tier model in the Gemini 3.5 generation, announced at Google I/O 2026. GA at launch and set as the default model in the Gemini app and AI Mode in Google Search. Architecture is not publicly disclosed. Vendor-reported benchmarks: GPQA Diamond 90.4, MMMU-Pro 81.2, SWE-bench Verified 78. Independent (Artificial Analysis, ~5 days post-launch): HLE 40.2 (exact match to vendor figure), Terminal-Bench 2.1 76.2, AA Intelligence Index v4.0: 55. Distinct from 'Gemini Omni', a separate video/world model announced at the same event.
Key Capabilities
- Default model in Gemini app and Search AI Mode
- GA at launch (2026-05-19)
- Multimodal input
Innovations
- ★Output speed 207.9 tok/s — AA Speed rank #2/148 at launch (Speed is the product identity behind the Flash name)
ERNIE 5.1
2026-05-08MoE flagship derived from ERNIE 5.0 via 'multi-dimensional elastic pre-training' as an optimal sub-network. Parameter counts are estimates (~1/3 of 5.0's total, ~1/2 of 5.0's active params); Baidu has not published exact figures. Cloud-only via Baidu Qianfan, no Western public API. Multimodal (text/image/audio/video) inherited from 5.0. Baidu claims ~6% of comparable training cost — unverified by third parties. No standalone technical report published. Self-reported benchmarks; non-EU hosting.
Key Capabilities
- Multimodal (text/image/audio/video)
- 128k context window
- MoE sub-network of ERNIE 5.0
- Cloud-only (Baidu Qianfan)
- Closed weights
Innovations
- ★Multi-Dimensional Elastic Pre-Training (sub-network extraction from ERNIE 5.0)
- ★~6% of comparable training cost (Baidu claim, unverified)
GPT-5.5 Instant
2026-05-05Default ChatGPT model, replacing GPT-5.3 Instant. API identifier 'chat-latest'. Hallucination-focused update with material gains on math and multimodal reasoning vs. its predecessor. Architecture not publicly disclosed.
Key Capabilities
- Default ChatGPT model
- API identifier: chat-latest
- Multimodal (text + image)
- Reduced hallucinations on high-stakes prompts
- Tool use / function calling
Innovations
- ★ChatGPT default for hundreds of millions of users
- ★52.5% Hallucination Reduction vs 5.3 Instant (high-stakes, internal eval)
- ★AIME 2025 Jump 65.4 to 81.2 vs 5.3 Instant
- ★Low-latency 'Instant' tier — OpenAI states GPT-5.5 matches GPT-5.4 per-token latency in real-world serving (no AA Instant entry yet)
Mistral Medium 3.5
2026-04-30128B dense flagship with 256k context, multimodal (text + vision), open weights under a modified MIT license. First Mistral flagship trained on the new 13,800-GPU Paris facility. Merges the Magistral reasoning line and the Devstral 2 coding line into one model. Mistral did not publish GPQA/HLE/MMLU-Pro at launch; granular third-party scores still pending. Independent composite available: AA Intelligence Index v4.0 = 39 (#2 in class), with very verbose output (~90M tokens vs. 16M class average).
Key Capabilities
- 128B dense parameters
- 256k context window
- Multimodal (text + vision)
- Reasoning + coding unified
- Open weights (Modified MIT)
- Self-hostable on 4 GPUs (~70GB VRAM Q4)
Innovations
- ★First flagship trained on Mistral's 13,800-GPU Paris facility
- ★Unified Magistral (reasoning) + Devstral 2 (coding) in one model
- ★Modified MIT release at flagship scale
Grok 4.3
2026-04-30Reasoning-first flagship with a 1M-token context window and native video input. Agentic-focused release at lower pricing than Grok 4.1.
Key Capabilities
- 1M context window
- Native video input
- Reasoning-first
Innovations
MiMo-V2.5-Pro
2026-04-27Xiaomi's flagship coding/agentic open-weights model. 1.02T total / 42B active hybrid MoE (70 layers, 384 routed experts, 8 active per token) with interleaved Sliding Window + Global Attention at a 6:1 ratio and 128-token window for long-context efficiency. 1M context, MIT licensed, day-0 SGLang/vLLM support. Reports SWE-bench Pro 57.2, GDPVal-AA Elo 1581, ClawEval 64% Pass³ at ~70K tokens per trajectory, and a SysY compiler-in-Rust task solved 233/233 in 4.3h with 672 tool calls. Project lead Fuli Luo (ex-DeepSeek).
Key Capabilities
- 1.02T total / 42B active (hybrid MoE)
- 1M context window
- Hybrid SWA + Global Attention (6:1 ratio, 128-token window)
- FP8 E4M3 mixed precision
- Day-0 SGLang/vLLM support
- MIT license
Innovations
- ★Interleaved Sliding Window + Global Attention (6:1)
- ★Lightweight MTP modules with dense FFNs
- ★Long-horizon agent trajectories (~70K tokens / 672 tool calls)
DeepSeek-V4-Pro
2026-04-24Preview release, available via API and open weights. 1.6T total / 49B active MoE with hybrid attention and manifold-constrained hyper-connections. Reports 27% of single-token inference FLOPs and 10% of KV cache vs DeepSeek-V3.2 in a 1M-token setting.
Key Capabilities
- 1.6T total / 49B active MoE
- 1M token context window
- Three reasoning effort modes (incl. Think Max)
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens
DeepSeek-V4-Flash
2026-04-24Preview release, available via API and open weights; superseded by DeepSeek-V4-Flash-0731 on 2026-07-31. Smaller 284B total / 13B active MoE sibling of V4-Pro, sharing hybrid attention and hyper-connection architecture.
Key Capabilities
- 284B total / 13B active MoE
- 1M token context window
- Three reasoning effort modes (incl. Think Max)
- MIT license
Innovations
- ★Hybrid attention: CSA + HCA
- ★Manifold-Constrained Hyper-Connections
- ★Shared architecture with V4-Pro at smaller scale
GPT-5.5
2026-04-23Agentic-focused flagship with 1M context window (400K in Codex). Thinking and Pro variants available. Codename 'Spud'.
Key Capabilities
- 1M context window
- Thinking mode
- Pro variant ($30/$180 per M tokens)
- 400K context in Codex
- Agentic workflows
- API pricing: $5/$30 per M tokens (base)
Innovations
- ★1M Context Window (largest OpenAI flagship)
- ★Codex-Optimized 400K Context
- ★Thinking + Pro Variant Lineup
ChatGPT Images 2.0
2026-04-21First OpenAI image model with reasoning capabilities. 2K resolution, multilingual text rendering, up to 8 images per prompt, web search integration. API name 'gpt-image-2'. Replaces DALL-E 3 (deprecated May 12, 2026).
Key Capabilities
- Text-to-Image
- 2K resolution output
- Reasoning-guided generation
- Multilingual text rendering (Japanese, Korean, Hindi, Bengali)
- Up to 8 images per prompt
- Web search integration
- API: gpt-image-2 ($8/$30 per M tokens)
Innovations
- ★Reasoning-Guided Image Generation
- ★Multilingual Text Rendering
- ★Web Search Integration for Image Context
- ★DALL-E 3 Successor
GR00T N1.7
2026-04-17Open vision-language-action model for humanoid robots, released in early access under Apache 2.0 — the first fully commercially licensed model of the GR00T line. A vision-language backbone is paired with a diffusion transformer head that denoises continuous actions via flow matching. NVIDIA credits the generalization and language-following gains over N1.6 to 20,000 hours of EgoScale human egocentric video in pretraining. At 3B parameters it is small enough to run on local hardware, which is why it also appears in the Local tab.
Key Capabilities
- Vision-language-action robot control
- Flow-matching action transformer head
- Humanoid whole-task policies
- Apache 2.0 open weights
Innovations
- ★Open, commercially licensed humanoid VLA
- ★20K hours of EgoScale human video in pretraining
Claude Opus 4.7
2026-04-16Anthropic's most capable generally available model. Default in Claude Code. Dense architecture with new tokenizer, file-system-based cross-session memory, new 'xhigh' effort level, and task budgets (public beta) for agentic loops. 1M context window with no long-context premium. Sits below Mythos Preview in capability, above Opus 4.6. $5/M input, $25/M output tokens.
Key Capabilities
- 1M context window
- Vision input up to 3.75 MP
- Computer use with 1:1 pixel mapping
- Agentic workflows
- Cross-session file-system memory
Innovations
- ★xhigh Effort Level
- ★Task Budgets (Public Beta)
- ★File-System Memory
- ★New Tokenizer
Qwen3.6-35B-A3B
2026-04-16Successor family to Qwen3.5 from the Qwen team, focused on agentic coding. 35B total / 3B active MoE with 262K context window and thinking preservation across conversation turns. Compatible with OpenClaw, Claude Code, and Cline. Apache 2.0 weights on HuggingFace.
Key Capabilities
- 35B total / 3B active MoE
- 262K context window
- Agentic coding focus
- Repository-level reasoning
- Thinking preservation across turns
- Apache 2.0
Innovations
- ★Thinking Preservation Across Conversation
- ★Repository-Level Reasoning
- ★OpenClaw/Claude Code/Cline Compatible
π0.7
2026-04-16Steerable robot foundation model from Physical Intelligence, presented as the point where VLAs start generalising rather than replaying training data. Physical Intelligence reports 82.1% success on trained tasks and 47.3% zero-shot on unseen ones across seven robot embodiments and 50+ manipulation tasks, and demonstrates the model operating an air fryer it had only encountered in two fragmentary training episodes. Weights are not public.
Key Capabilities
- Vision-language-action robot control
- Zero-shot spatial generalization
- Semantic steering from language alone
- Multi-step compositional manipulation
Innovations
- ★Emergent task generalization beyond the training distribution
- ★One policy across seven robot embodiments
Muse Spark
2026-04-08First model from Meta Superintelligence Labs under Alexandr Wang. Powers Meta AI across Facebook, Instagram, WhatsApp, Messenger and Ray-Ban glasses. Natively multimodal (text/image/voice input, text output) with multi-agent 'Contemplating mode'. Closed-source departure from the Llama lineage. Free on meta.ai and the Meta AI app; API in private preview. Consumer-focused, strong in health and multimodal; coding gap acknowledged by Meta.
Key Capabilities
- Multimodal perception
- Voice input
- Reasoning and agentic tasks
- Powers Meta AI products
- Ray-Ban smart glasses integration
- Health Q&A
Innovations
- ★Meta Superintelligence Labs Debut
- ★Contemplating Mode (multi-agent)
- ★Closed-source Departure from Llama
Claude Mythos Preview
2026-04-07Anthropic's 'Capybara tier' flagship, withheld from public release due to cybersecurity risk. Available only to 12 Project Glasswing enterprise partners (incl. Amazon, Apple, Google, Microsoft, NVIDIA) via a $100M credit pool. No public API. $25/M input, $125/M output tokens.
Key Capabilities
- Frontier reasoning
- Advanced coding
- Agentic workflows
- Mathematical reasoning
Innovations
- ★Capybara Tier
- ★Project Glasswing Restricted Access
- ★Withheld for Cybersecurity Risk
Happy Horse 1.0
2026-04-07Video generation model from Alibaba's ATH AI Innovation Unit. 15B-parameter single-stream Transformer (40 layers) with native audio/video/text token fusion. #1 in both text-to-video and image-to-video on Artificial Analysis Video Arena (Elo blind test). Initially released anonymously; confirmed as Alibaba on April 10, 2026. Beta — no public API or released weights.
Key Capabilities
- Text-to-video
- Image-to-video
- Native audio/video/text token fusion
Innovations
- ★Single-stream 40-layer Transformer
- ★Video Arena #1 (T2V and I2V)
- ★Native Multimodal Token Fusion
Qwen3.6 Plus
2026-03-31Flagship cloud API model with a 1M-token context window and always-on chain-of-thought reasoning. Agentic coding focus.
Key Capabilities
- 1M context window
- Always-on chain-of-thought
- Agentic coding
Innovations
Midjourney V8 Alpha
2026-03-17Major architectural overhaul on rewritten codebase. ~5x faster generation than V7, native 2K resolution, significantly improved text rendering, better prompt adherence for complex multi-element compositions.
Key Capabilities
- ~5x faster generation
- Native 2K resolution (--hd)
- Improved text rendering
- Better prompt adherence
- New --q 4 quality mode
Innovations
- ★Rewritten Codebase
- ★Native 2K Resolution
- ★Enhanced Text Rendering
- ★Scene Coherence Mode
GPT-5.4 mini
2026-03-17Fast reasoning model approaching GPT-5.4 performance at 2x the speed, optimized for sub-agent workloads.
Key Capabilities
- Reasoning (configurable effort)
- Multimodal (text + image)
- Tool use / function calling
- Computer use
- 2x faster than GPT-5 mini
Innovations
- ★Sub-Agent Architecture Optimized
- ★Near-Flagship Performance at Mini Scale
- ★2x Speed vs GPT-5 mini
GPT-5.4 nano
2026-03-17Smallest, cheapest GPT-5.4 variant for classification, extraction, and cost-sensitive sub-agent tasks.
Key Capabilities
- Multimodal (text + image)
- Tool use / function calling
- Classification & extraction
- Cost-optimized ($0.20/1M input)
Innovations
- ★Ultra Cost-Efficient Reasoning
- ★Sub-Agent Era Design
- ★Smallest GPT-5.4 Variant
Mistral Small 4
2026-03-16119B MoE model (6B active) unifying reasoning, vision, and coding with configurable reasoning effort.
Key Capabilities
- 119B MoE (6B active)
- Configurable reasoning effort
- Multimodal (vision)
- Agentic coding
- 256k context
- Apache 2.0
Innovations
- ★Unified Magistral + Pixtral + Devstral
- ★128-Expert MoE Architecture
- ★40% Latency Reduction vs Small 3
- ★3x Throughput vs Small 3
GPT-5.3 Instant
2026-03-05Fast everyday model with reduced hallucinations and improved web search.
Key Capabilities
- Reduced hallucinations
- Improved web search
- Natural conversation
- Creative writing
Innovations
- ★26.8% Hallucination Reduction (web)
- ★19.7% Hallucination Reduction (internal)
- ★Balanced Reasoning + Search
GPT-5.4 Thinking
2026-03-05Multi-step reasoning model with native computer-use and 1M token context.
Key Capabilities
- Multi-step reasoning
- Native computer use
- 1M token context (API)
- Tool-heavy workflows
- Agentic tasks
Innovations
- ★Native Computer Use
- ★Tool Search
- ★Efficient Token Usage
- ★28-point OSWorld Jump
GPT-5.4 Pro
2026-03-05Highest-capability model optimized for quality and depth over speed.
Key Capabilities
- Highest capability
- Decision-ready outputs
- Extended reasoning
- Native computer use
- 1M token context (API)
Innovations
- ★ARC-AGI-2 Leader (83.3%)
- ★Quality over Speed Optimization
Nano Banana 2
2026-02-26Pro-quality AI image generation at Flash speed, combining Nano Banana Pro features with Gemini Flash performance.
Key Capabilities
- 4K image generation
- Text rendering in images
- Real-time knowledge integration
- Multi-character consistency
- Extreme aspect ratios
Innovations
- ★Gemini 3.1 Flash Image
- ★Web Search Integration
- ★Multi-language Text Rendering
- ★Subject Consistency (up to 5 characters)
Gemini 3.1 Flash Image
2026-02-26Google DeepMind's Nano Banana 2 (API gemini-3.1-flash-image-preview), third in the Nano Banana family after the original (Gemini 2.5 Flash, Aug 2025) and Nano Banana Pro (Gemini 3 Pro, Nov 2025). Brings Pro-tier image quality to the Flash architecture at roughly half the per-image API price (~$0.067 vs $0.134 for a 2K image).
Key Capabilities
- Text-to-Image
- Pro-quality at Flash speed
- Strong text rendering
- Lower per-image cost
Innovations
- ★Pro Quality at Flash Speed
- ★Half-Price 2K Images
Gemini 3.1 Pro
2026-02-19Major reasoning upgrade with record GPQA and 2x ARC-AGI-2 improvement over Gemini 3 Pro.
Key Capabilities
- 1M context window
- 64K output
- Natively multimodal
- Agentic workflows
Innovations
- ★ARC-AGI-2 Regime Change
- ★Deep Think Reasoning
- ★SVG Animation Generation
Claude Sonnet 4.6
2026-02-17Near-flagship performance at mid-tier pricing with 1M context window and advanced computer use.
Key Capabilities
- 1M context window
- Advanced computer use
- Adaptive thinking
- Financial analysis
Innovations
- ★Flagship-tier at Mid-tier Pricing
- ★ARC-AGI-2 Leap
- ★OSWorld Computer Use
Qwen3.5
2026-02-16Major architectural upgrade over Qwen3. 397B MoE (17B active) flagship with 1M context, natively multimodal, 8-19x higher decoding throughput. Covers 200+ languages.
Key Capabilities
- 397B total / 17B active MoE (flagship)
- 1M context window
- Natively multimodal
- Reasoning by default
- 200+ languages
- Apache 2.0
Innovations
- ★8-19x Throughput vs Qwen3-Max
- ★Native Multimodal All Sizes
- ★122B-A10B Runs on MacBook 64GB
Doubao Seed 2.0 Pro
2026-02-14Frontier-Flaggschiff der Seed-2.0-Familie (Pro/Lite/Mini/Code) von ByteDance, 256k Kontext, Schwerpunkt Reasoning, Coding und Agent-Tasks. Treibt die Doubao-App, Chinas meistgenutzten KI-Chatbot. Proprietär über Volcano Engine.
Key Capabilities
- 256k Context
- Reasoning / Agent
- Coding
- Multilingual
Innovations
- ★Seed-2.0-Foundation-Familie
- ★Volcano-Engine-API
Seedance 2.0
2026-02-12ByteDance text/image/audio-to-video model with unified audio-video joint generation. Generates 4-15s clips up to 1080p across multiple aspect ratios. Available via Volcano Engine and API.
Key Capabilities
- Text-to-video
- Image-to-video
- Audio-video joint generation
- Up to 1080p output
- 4-15s clips
Innovations
- ★Unified multimodal audio-video architecture
- ★Multimodal reference inputs (text/image/audio/video)
- ★Director-level camera & lighting control
GLM-5
2026-02-11744B MoE (40B active) open-weight model with DeepSeek Sparse Attention and strong agentic/frontend coding performance.
Key Capabilities
- 744B total / 40B active MoE
- 200K context
- Agentic workflows
- Frontend coding (98% build success)
- MIT license
Innovations
- ★DeepSeek Sparse Attention
- ★Slime Async RL Framework
- ★98% Frontend Build Success Rate
- ★Stealth-Launched as Pony Alpha on OpenRouter
GPT-5.3 Codex
2026-02-05Agentic coding model, 25% faster, first model instrumental in creating itself.
Key Capabilities
- Agentic coding
- 25% faster
- Self-developing
- Tool use
Innovations
- ★Self-Development
- ★Terminal-Bench SOTA
- ★OSWorld Leader
Claude Opus 4.6
2026-02-05Next-generation flagship model with adaptive thinking and agent teams.
Key Capabilities
- Adaptive thinking
- 1M context window
- Agent teams
- 128K output
Innovations
- ★Agent Teams
- ★Adaptive Reasoning
- ★Claude Cowork
Kling 3.0
2026-02-05Latest flagship with AI Director capability and multi-shot storyboard.
Key Capabilities
- 15s video
- Native audio
- 2K/4K images
- AI Director
- Video generation
Innovations
- ★AI Director
- ★Multi-shot Storyboard
- ★Character/Voice Replication
Grok Imagine 1.0
2026-01-28High-definition text-to-video model with synchronized cinematic audio.
Key Capabilities
- Text-to-video
- Native audio
- Video editing
- HD output
Innovations
- ★End-to-end video+audio generation
- ★#1 Artificial Analysis video ranking
- ★Cinematic quality
Qwen3-Max-Thinking
2026-01-27Trillion-parameter reasoning model with adaptive tool use and test-time scaling.
Key Capabilities
- Advanced reasoning
- Adaptive tool use
- Test-time scaling
- Code interpretation
- Web search integration
- Video understanding
Innovations
- ★Reasoning Mode
- ★Adaptive Tool Invocation
- ★1T+ MoE Architecture
- ★Dynamic Compute Allocation
Kimi K2.5
2026-01-271T MoE (32B active) open-weight model with native multimodal vision and agent swarm support for up to 100 sub-agents.
Key Capabilities
- 1T total / 32B active MoE
- 262K context
- Natively multimodal (MoonViT 400M)
- Agent swarms (100 sub-agents)
- 1,500 parallel tool calls
- Modified MIT license
Innovations
- ★MoonViT Vision Encoder (400M params)
- ★100 Sub-Agent Swarm Support
- ★1,500 Parallel Tool Calls
- ★Top-Ranked Artificial Analysis & LMArena at Launch
ERNIE 5.0
2026-01-222.4T-parameter ultra-sparse MoE flagship in public preview via ERNIE Bot. Natively omnimodal (text/image/audio/video jointly modeled from pretraining). #1 Chinese model on LMArena Text and Vision (#8 globally). China-market only; no Western public API.
Key Capabilities
- Text
- Image
- Audio
- Video
- Omnimodal reasoning
- Agentic workflows
Innovations
- ★Natively Omnimodal Unified Autoregressive Framework
- ★Ultra-Sparse MoE (<3% activation per inference)
- ★PaddlePaddle Framework
Gemini 3 Flash
2025-12-17Flash-tier model offering frontier-level reasoning at lower cost and higher speed. Predecessor to Gemini 3.5 Flash.
Key Capabilities
- High-speed reasoning
- Multimodal input
Innovations
GPT Image 1.5
2025-12-16OpenAI's second-generation flagship image model (API gpt-image-1.5), rolled out globally to all ChatGPT tiers. About 4x faster generation, stronger instruction adherence, edit precision that preserves lighting and composition, and improved text rendering for dense typography. The high-fidelity variant is the strongest of the GPT Image 1.x line on the LMArena Text-to-Image arena.
Key Capabilities
- Text-to-Image
- High-fidelity generation
- ~4x faster generation
- Instruction-preserving edits
- Dense text rendering
Innovations
- ★High-Fidelity Editing
- ★4x Faster Generation
GWM-1
2025-12-11Runway's first General World Model: an autoregressive model built on top of Gen-4.5 that generates frame by frame, runs in real time at 24 fps and is steered interactively through camera pose, robot commands or audio. Ships as three variants, GWM Worlds for explorable environments, GWM Avatars for speaking characters with facial expression and lip sync, and GWM Robotics for synthetic robot training data, which Runway plans to merge into one model. Announced together with native audio for Gen-4.5 and pitched as more general than Google's Genie 3: a simulator for training agents in domains such as robotics and life sciences.
Key Capabilities
- Real-time frame-by-frame world simulation at 24 fps
- Interactive control via camera pose, robot commands and audio
- Explorable worlds, avatars and robotics variants
- Synthetic training data for robots
Innovations
- ★Autoregressive world model built on a video generation model (Gen-4.5)
- ★Action-conditioned real-time simulation
Seedream 4.5
2025-12-04All-round upgrade to Seedream 4.0 with stronger in-image text rendering, roughly 10× faster generation than 4.0, and multi-reference fusion accepting up to 10 reference images. Output up to 2048×2048, offered as text-to-image and image-edit variants.
Key Capabilities
- Text-to-Image
- Image editing
- Strong in-image text rendering
- Multi-reference fusion (up to 10 images)
- Up to 2048×2048 output
Innovations
- ★~10× faster than Seedream 4.0
- ★Multi-reference fusion
- ★Improved text rendering
Kling 2.6
2025-12-03First simultaneous audio-visual generation model.
Key Capabilities
- Audio-visual generation
- Speech & dialogue
- 1080p/10s
- Video generation
Innovations
- ★Simultaneous Audio-Visual
- ★Multi-language Audio
- ★Ambient Sound Generation
Mistral 3 Family
2025-12-02Flagship Mistral Large 3 (675B MoE) and efficient Ministral 3 edge models.
Key Capabilities
- 675B MoE
- Multimodal
- 256k Context
- Edge-optimized Ministral
Innovations
- ★Native Multimodal MoE
- ★Edge-Cloud Model Family
- ★Apache 2.0 Weights
DeepSeek-V3.2
2025-12-01Successor to DeepSeek V3. 671B MoE (37B active) with DeepSeek Sparse Attention for long-context efficiency and large-scale agentic tool-use pipeline.
Key Capabilities
- 671B total / 37B active MoE
- 164K context
- DeepSeek Sparse Attention (DSA)
- Agentic tool-use (1,800+ environments)
- Strong reasoning without thinking mode
- MIT license
Innovations
- ★DeepSeek Sparse Attention (DSA)
- ★Large-Scale Task Synthesis Pipeline
- ★1,800+ Agentic Environments
Gen-4.5
2025-12-01Advanced AI video generation model with physics-based realism and director controls.
Key Capabilities
- 1080p Cinematic Video
- Physics Simulation
- Director Mode 2.0
- Motion Brush 3.0
Innovations
- ★General World Model
- ★Physics-based Realism
- ★Latent Diffusion Transformers
Claude Opus 4.5
2025-11-24Next-generation flagship model with breakthrough reasoning and research capabilities.
Key Capabilities
- Advanced reasoning
- Scientific research
- Long-context understanding
- Multi-step planning
Innovations
- ★Enhanced Constitutional AI
- ★Extended Context Processing
- ★Advanced Tool Integration
Gemini 3 Pro Image Preview
2025-11-20Advanced image generation model building upon Nano Banana.
Key Capabilities
- Text-to-Image
- High fidelity generation
- Multimodal integration
Innovations
- ★Nano Banana Enhancement
- ★Advanced Image Synthesis
- ★Gemini 3 Integration
GPT-5.1-Codex-Max
2025-11-19Specialized ultra-high-performance coding model with breakthrough architecture.
Key Capabilities
- Advanced code generation
- Multi-language mastery
- System-level reasoning
Innovations
- ★Neural Code Synthesis
- ★Real-time Code Optimization
- ★Integrated Debugging
ChatGPT 5.1
2025-11-12Major update with "Instant" and "Thinking" modes for adaptive reasoning.
Key Capabilities
- Adaptive reasoning
- Instant/Thinking modes
- Personalization
Innovations
- ★Node-based Agent Builder
Marble
2025-11-12World Labs' first commercial product and, in its own words, a frontier multimodal world model. Text prompts, photos, videos, panoramas or coarse 3D layouts become persistent, editable 3D environments that export as Gaussian splats, meshes or video. The distinction from frame-by-frame world generators is that the scene geometry is fixed once generated: a world can be revisited, edited and walked through again instead of being re-dreamed on every move. Available to everyone from launch in free and paid tiers; Marble 1.1 later added auto-expanding worlds, and the World API followed in January 2026.
Key Capabilities
- Text, image, video and panorama to 3D world
- Persistent, editable 3D environments
- Export as Gaussian splats, meshes or video
- Free and paid tiers
Innovations
- ★Persistent world geometry instead of per-frame regeneration
- ★First commercial world model product
Odyssey-2
2025-10-27Odyssey's general-purpose world model, trained on a large corpus of general video and interaction data rather than on game engines or synthetic environments, which the company presents as evidence that basic physics, dynamics and behaviours can be learned from video alone. It generates interactive video that responds to typed input within about fifty milliseconds and keeps running for minutes rather than seconds; Odyssey describes the interaction like using a language model: you type, and the video responds. Publicly usable at experience.odyssey.ml, later extended by Odyssey-2 Pro (API access for embedding continuous simulations into applications) and Odyssey-2 Max.
Key Capabilities
- Interactive video generation from typed input
- Response to interaction in about fifty milliseconds
- Multi-minute simulations
- API access via Odyssey-2 Pro
Innovations
- ★World model trained on general video and interaction data rather than synthetic environments
Claude Haiku 4.5
2025-10-01Ultra-fast, edge-capable model.
Key Capabilities
- On-device potential
- Sub-10ms latency
- Cost efficiency
Innovations
Gemini Robotics 1.5
2025-09-25Vision-language-action model from Google DeepMind that turns camera input and natural-language instructions into robot motor commands. Shipped as a pair: Gemini Robotics 1.5 executes, while Gemini Robotics-ER 1.5 does the embodied reasoning — it can call digital tools such as web search to plan a task before handing execution over. ER 1.5 is available to developers through the Gemini API in Google AI Studio; the VLA itself went to selected partners only.
Key Capabilities
- Vision-language-action robot control
- Embodied reasoning with tool use (ER 1.5)
- Motion transfer across robot embodiments
- Thinks before acting
Innovations
- ★Split VLA / embodied-reasoning architecture
- ★Web search as a planning step for physical tasks
Kling 2.5 Turbo
2025-09-23Faster, cheaper generation with improved motion and style consistency.
Key Capabilities
- 40% faster
- High-motion sequences
- Style consistency
- Video generation
Innovations
- ★Turbo Mode
- ★30% Cost Reduction
- ★Real-world Physics
DeepSeek-Terminus
2025-09-22Experimental model pushing the boundaries of open weights.
Key Capabilities
- Experimental arch
- Community research
- Novel attention
Innovations
GPT-5-codex
2025-09-15Specialized coding model with architectural breakthroughs.
Key Capabilities
- Architectural breakthrough
- Self-improving code
- System design
Innovations
- ★Integration into VS-Code
Seedream 4.0
2025-09-09Next-generation multimodal image model that handles generation and instruction-based editing in one system, with high-definition output up to 4K.
Key Capabilities
- Text-to-Image
- Instruction-based image editing
- Multimodal image tasks
- Up to 4K output
Innovations
- ★Unified generation + editing
- ★Up to 4K resolution
Claude Sonnet 4.5
2025-09-01Balanced model with next-gen capabilities.
Key Capabilities
- Real-time collaboration
- Voice mode
- Advanced tool use
Innovations
Grok Code Fast
2025-08-28Specialized high-speed coding model.
Key Capabilities
- Instant code gen
- Repo-level understanding
- Test driven dev
Innovations
DeepSeek-V3.1
2025-08-21Incremental update with better instruction following.
Key Capabilities
- Instruction following
- Long context
- Math improvements
Innovations
Claude Opus 4.1
2025-08-05Massive reasoning model for complex R&D tasks.
Key Capabilities
- Deep research
- Scientific discovery
- Long-horizon planning
Innovations
Genie 3
2025-08-05Google DeepMind's first real-time interactive general-purpose world model. From a text prompt it generates a navigable environment at 24 frames per second and 720p that stays consistent for a few minutes, and the world can be changed mid-simulation with further text prompts, which DeepMind calls promptable world events. Released as a limited research preview and positioned as a training environment for embodied agents rather than as a video generator; a public demo followed in January 2026 as Project Genie.
Key Capabilities
- Real-time interactive world generation
- Text-prompted world events during simulation
- Multi-minute visual consistency
- Agent training environments
Innovations
- ★First real-time interactive general-purpose world model
- ★Promptable world events
Qwen3-Coder
2025-07-15Code-specialized model optimized for software engineering.
Key Capabilities
- Code generation
- Agentic coding
- Multi-file editing
- Tool use
Innovations
- ★Code-specialized Training
- ★Agentic SWE Pipeline
- ★Multi-file Context
Gemini 2.5
2025-06-17Iterative update with enhanced speed and reasoning.
Key Capabilities
- Flash/Pro variants
- Lower latency
- Higher accuracy
Innovations
- ★Nano Banana image generation
Seedance 1.0
2025-06-11ByteDance's first Seedance video-generation model, introduced at the Volcano Engine Force conference. Fast text- and image-to-video generation with multi-shot narrative sequences.
Key Capabilities
- Text-to-video
- Image-to-video
- Multi-shot narratives
- Fast generation
Innovations
- ★Multi-shot narrative video generation
- ★First Seedance video model
Magistral Medium
2025-06-10Mistral's first reasoning model. API and enterprise deployment only — no published weights. The 24B Magistral Small released the same day carries open weights and is tracked separately in the local catalog.
Key Capabilities
- Enterprise focus
- Privacy first
- Specialized domains
Innovations
Claude Sonnet 4
2025-05-22Next-generation Sonnet with improved instruction following and coding.
Key Capabilities
- Improved instruction following
- Enhanced coding
- Better reasoning
Innovations
- ★Improved Alignment
- ★Better Calibration
Claude Opus 4
2025-05-22Frontier reasoning model for complex tasks and sustained autonomous work.
Key Capabilities
- Deep research
- Sustained autonomy
- Complex reasoning
Innovations
- ★Extended Autonomy
- ★Deep Analysis
- ★Multi-hour Tasks
o3
2025-04-16Full reasoning model with significant improvements over o1.
Key Capabilities
- Advanced reasoning
- Tool use
- Agentic coding
Innovations
- ★Improved Reasoning Chain
- ★Tool Integration
- ★Multi-step Planning
Seedream 3.0
2025-04-16ByteDance's bilingual (Chinese/English) text-to-image foundation model. Native 2048×2048 output with fast inference (~3s for a 1K image on an A100). Available via the Doubao and Jimeng apps.
Key Capabilities
- Text-to-Image
- Bilingual (Chinese/English) prompts
- Native 2048×2048 output
- Fast inference (~3s at 1K)
Innovations
- ★Bilingual text-to-image foundation model
- ★Native 2K output
Kling 2.0
2025-04-15Cinematic quality video with advanced physics and multi-element editing.
Key Capabilities
- Cinematic 1080p
- Physics simulation
- Multi-element editor
- Video generation
Innovations
- ★Multi-Modal Visual Language
- ★Multi-Element Editor
- ★Advanced Physics
Llama 4 Scout
2025-04-05109B MoE (17B active) multimodal model with 10M token context window. Outperforms Gemma 3, Gemini 2.0 Flash, and Mistral 3.1 Small.
Key Capabilities
- 109B total / 17B active MoE
- 10M context window
- Natively multimodal
- 16 experts, 1 active
- Open weights (Llama license)
Innovations
- ★Early Fusion Architecture
- ★10M Token Context (Longest Open-Weight)
- ★Interleaved Attention / MoE Blocks
Llama 4 Maverick
2025-04-05400B MoE (17B active) multimodal model. Matches Gemini 2.5 Pro and DeepSeek-V3 on reasoning benchmarks at significantly lower serving cost.
Key Capabilities
- 400B total / 17B active MoE
- 128 experts, 1 active
- Natively multimodal
- 1M context window
- Open weights (Llama license)
Innovations
- ★128-Expert MoE (Largest Open-Weight MoE)
- ★Early Fusion Architecture
- ★Frontier Performance at Low Active Params
Llama 4 Behemoth
2025-04-052T MoE (288B active) teacher model. Announced alongside Scout/Maverick, still in training at launch. Tops benchmarks in STEM reasoning.
Key Capabilities
- 2T total / 288B active MoE
- 16 experts, 4 active
- STEM reasoning
- Natively multimodal
- In training at announcement
Innovations
- ★Largest Llama Model Ever
- ★Teacher Model Architecture
- ★288B Active Parameters
Midjourney V7
2025-04-03Stunning precision with improved text and image handling.
Key Capabilities
- Draft Mode
- Omni Reference
- Rich textures
- Better hands/bodies
Innovations
- ★Draft Mode
- ★Omni Reference
- ★Enhanced Precision
Runway Gen-4
2025-04-01Next-gen model with improved character and environment consistency.
Key Capabilities
- Character consistency
- Environment persistence
- 16s videos
- Gen-4 Turbo
Innovations
- ★Consistency Engine
- ★Extended Duration
- ★Turbo Mode
Gemini 2.5 Pro
2025-03-25Thinking model with enhanced reasoning capabilities.
Key Capabilities
- Thinking/reasoning
- Code generation
- Multimodal
Innovations
- ★Thinking Mode
- ★Enhanced Reasoning
- ★Experimental Preview
Hunyuan-T1
2025-03-21Tencent's deep-thinking reasoning flagship, built on the TurboS Hybrid-Transformer-Mamba MoE base and post-trained with large-scale curriculum reinforcement learning. First Hunyuan entry in the frontier reasoning race.
Key Capabilities
- Deep-thinking chain-of-thought reasoning
- Hybrid-Transformer-Mamba MoE base (TurboS)
- Fast decoding on long inputs
Innovations
- ★First ultra-large Mamba-hybrid reasoning model
- ★Curriculum reinforcement-learning post-training
Command A
2025-03-13111B Dense Enterprise-Flagship (command-a-03-2025), 256k Kontext, 23 Sprachen, Fokus auf Tool-Use, agentische Workflows und Translation. Läuft auf wenigen GPUs, offene Gewichte. Direkter Vorgänger von Command A+. Vectara HHEM 9.3% Halluzination.
Key Capabilities
- 111B Dense
- Enterprise / Tool-Use / Agentic
- 23 Sprachen
- 256k Context
- Open weights
Innovations
- ★Effizientes Enterprise-Flagship
- ★Open weights
Claude 3.7 Sonnet
2025-02-24Hybrid model with extended thinking for complex reasoning tasks.
Key Capabilities
- Extended thinking
- Enhanced coding
- Agentic reliability
Innovations
- ★Extended Thinking Mode
- ★Hybrid Reasoning
- ★Improved Tool Use
Grok-3 Mini
2025-02-17Lightweight reasoning model from xAI. Faster and cheaper than Grok-3 with configurable thinking for speed/quality tradeoffs.
Key Capabilities
- Lightweight reasoning
- Configurable thinking mode
- Faster inference than Grok-3
- Cost-optimized
Innovations
- ★Thinking Mode (Low/High)
- ★Reasoning at Mini Scale
Gemini 2.0 Flash
2025-02-05Next-gen Flash model with native tool use and multimodal generation.
Key Capabilities
- Native tool use
- Multimodal generation
- Fast inference
Innovations
- ★Native Tool Use
- ★Multimodal Output
- ★Agentic Capabilities
o3-mini
2025-01-31Smaller, faster reasoning model bringing o-series capabilities to a more accessible form.
Key Capabilities
- Reasoning
- Fast inference
- Cost-effective
Innovations
- ★Compact Reasoning
- ★Efficient CoT
- ★Accessible Intelligence
Mistral Small 3
2025-01-3024B open-weight model under Apache 2.0 with strong efficiency.
Key Capabilities
- 24B Parameters
- Apache 2.0
- On-device capable
Innovations
- ★Efficient 24B Scale
- ★Apache 2.0 License
- ★Edge Deployment
DeepSeek-R1
2025-01-20Reasoning model capable of self-verification.
Key Capabilities
- Chain of thought
- Self-verification
- RL training
Innovations
- ★Pure RL Reasoning
- ★Self-verification
- ★Emergent CoT Behavior
DeepSeek-V3
2024-12-25Latest iteration with improved benchmarks.
Key Capabilities
- SOTA Open Weights
- FP8 Training
- Multi-token prediction
Innovations
- ★FP8 Training
- ★Multi-token Prediction
- ★Auxiliary Loss-Free Load Balancing
Kling 1.6
2024-12-19Enhanced realism with better motion and prompt understanding.
Key Capabilities
- Enhanced realism
- Better expressions
- Video generation
- Style preservation
Innovations
- ★DeepSeek Prompt Understanding
- ★Elements Feature
- ★195% Quality Boost
Llama 3.3
2024-12-06Efficient 70B text model matching Llama 3.1 405B on key benchmarks.
Key Capabilities
- 70B Parameters
- 405B-level performance
- Cost efficient
Innovations
- ★Efficiency Breakthrough
- ★Distillation Techniques
QwQ-32B-Preview
2024-11-27Reasoning-focused model with chain-of-thought capabilities.
Key Capabilities
- 32B Parameters
- Chain-of-thought
- Math reasoning
- Self-reflection
Innovations
- ★Deliberative Reasoning
- ★Extended Thinking
- ★Self-Correction
Flux.1 Tools
2024-11-21Suite of editing tools including Fill, Depth, Canny, and Redux.
Key Capabilities
- Inpainting
- Outpainting
- Depth control
- Edge control
Innovations
- ★Flux.1 Fill
- ★Flux.1 Depth
- ★Flux.1 Canny
- ★Flux.1 Redux
Pixtral Large
2024-11-18124B multimodal model with frontier-class vision capabilities.
Key Capabilities
- 124B Parameters
- Frontier vision
- Multimodal
Innovations
- ★Large-scale Vision
- ★128k Context
- ★Multimodal Reasoning
Stable Diffusion 3.5 Medium
2024-10-29Compact 2.5B parameter variant optimized for consumer GPUs. Runs on hardware with 4-16 GB VRAM.
Key Capabilities
- 2.5B parameters
- Text-to-image
- Consumer GPU optimized (4-16 GB VRAM)
- Open weights (Stability Community License)
Innovations
- ★Efficient MMDiT at 2.5B Scale
- ★Consumer Hardware Target
Aya Expanse 32B
2024-10-2432B multilinguales Open-Weight-Modell aus der Aya-Forschung von Cohere Labs, 23 Sprachen, Schwerpunkt auf non-English-Evaluierungen (Global-MMLU, mArenaHard). Die englische Standard-Suite reportet der Vendor nicht, multilinguale Stärke und offene Gewichte sind der Fokus. Vectara HHEM 10.9% Halluzination.
Key Capabilities
- 32B Dense
- Multilingual (23 Sprachen)
- Open weights
- Aya-Forschungslinie
Innovations
- ★Multilingual-fokussiertes Open-Weight-Modell
- ★Data Arbitrage und multilinguales Preference-Training
Claude 3.5 Sonnet v2
2024-10-22Upgraded Sonnet with computer use capabilities.
Key Capabilities
- Computer use
- Enhanced coding
- Tool use
Innovations
- ★Computer Use
- ★Agent Capabilities
- ★Real-world Interaction
Claude 3.5 Haiku
2024-10-22Fastest model in its intelligence class.
Key Capabilities
- Low latency
- High intelligence
- Tool use
Innovations
Stable Diffusion 3.5 Large
2024-10-22Flagship 8.1B parameter open-weight image model with improved quality and prompt adherence. Available in standard and Turbo variants.
Key Capabilities
- 8.1B parameters
- Text-to-image
- Large + Large Turbo variants
- Open weights (Stability Community License)
Innovations
- ★MMDiT-X Architecture
- ★QK-Normalization
- ★Turbo Distillation Variant
Flux 1.1 Pro
2024-10-02Improved flagship model with enhanced quality and speed.
Key Capabilities
- Faster generation
- Better quality
- Professional use
Innovations
- ★Speed Improvements
- ★Quality Enhancement
- ★Pro Features
Kling 1.5
2024-09-19Improved quality with Motion Brush and camera controls.
Key Capabilities
- 1080p HD
- Motion Brush
- Camera controls
- Video generation
Innovations
- ★Motion Brush
- ★Camera Movement Controls
- ★95% Quality Improvement
Pixtral 12B
2024-09-17First multimodal model from Mistral with vision capabilities.
Key Capabilities
- Vision understanding
- 12B Parameters
- Open weights
Innovations
- ★Multimodal Vision
- ★Variable Image Resolution
- ★Native Vision Encoder
OpenAI o1
2024-09-12Reasoning model designed to spend more time thinking before responding.
Key Capabilities
- Chain of thought
- Complex math/coding
- Self-correction
Innovations
- ★Chain of Thought Reasoning
- ★Test-time Compute Scaling
- ★Self-correction via RL
FLUX.1 [dev]
2024-08-01Open-weight text-to-image model from Black Forest Labs. Guidance-distilled 12B param model for non-commercial use.
Key Capabilities
- 12B parameters
- Text-to-image
- Guidance distilled
- Open weights (non-commercial)
Innovations
- ★Flow Matching Architecture
- ★Hybrid Transformer (MMDiT)
- ★Guidance Distillation
FLUX.1 [schnell]
2024-08-01Fastest FLUX variant, Apache 2.0 licensed. Optimized for local inference and rapid prototyping.
Key Capabilities
- 12B parameters
- Text-to-image
- Fastest FLUX variant
- Apache 2.0 license
Innovations
- ★Timestep Distillation (4 steps)
- ★Apache 2.0 Open Weight Image Gen
Midjourney V6.1
2024-07-30Faster generation with more coherent images and precise details.
Key Capabilities
- 25% faster
- Coherent images
- Precise textures
Innovations
- ★Speed Optimization
- ★Detail Enhancement
- ★Texture Precision
Mistral Large 2
2024-07-24Significant improvements in code generation and reasoning.
Key Capabilities
- 123B Parameters
- Code generation
- Function calling
Innovations
- ★123B Parameters
- ★Code-focused Training
- ★Advanced Function Calling
Llama 3.1
2024-07-23Massive 405B open-source model competing with frontier closed models.
Key Capabilities
- 405B Parameters
- 128k Context
- Tool use
- Multilingual
Innovations
- ★405B Open Weights
- ★128k Context
- ★Synthetic Data Scaling
Mistral Nemo
2024-07-1812B parameter model developed in collaboration with NVIDIA.
Key Capabilities
- 12B Parameters
- Apache 2.0
- 128k context
Innovations
- ★NVIDIA Collaboration
- ★Tekken Tokenizer
- ★Efficient 12B Scale
ERNIE 4.0 Turbo
2024-06-28Speed-optimized variant with reduced latency.
Key Capabilities
- Faster inference
- Advanced reasoning
- Tool use
Innovations
- ★Optimized Inference
- ★Reduced Latency
- ★Cost Efficiency
Claude 3.5 Sonnet
2024-06-21Significant leap in coding and reasoning capabilities.
Key Capabilities
- SOTA Coding
- Artifacts UI
- Fast inference
Innovations
- ★Artifacts UI
- ★Real-time Collaboration
- ★Interactive Content Generation
Runway Gen-3 Alpha
2024-06-18Advanced video generation with creative control and motion tools.
Key Capabilities
- Creative control
- Motion brush
- Camera controls
- Director mode
Innovations
- ★Advanced Camera Controls
- ★Motion Brush
- ★Keyframing
Stable Diffusion 3
2024-06-12Next-gen architecture with Multimodal Diffusion Transformer.
Key Capabilities
- MMDiT architecture
- Improved text rendering
- Better composition
Innovations
- ★Multimodal Diffusion Transformer
- ★Flow Matching
- ★Triple Text Encoders
Kling 1.0
2024-06-06First consumer-accessible video generation model with up to 2 min output.
Key Capabilities
- Text-to-video
- 1080p/30fps
- Up to 2 min
- Complex motion
Innovations
- ★DiT Architecture
- ★3D VAE Compression
- ★Spatiotemporal Simulation
Gemini 1.5 Flash
2024-05-14Lightweight model optimized for speed and efficiency.
Key Capabilities
- Low latency
- High throughput
- Cost effective
Innovations
- ★Efficient MoE
- ★Low Latency Optimization
- ★Cost-effective Inference
DeepSeek-V2
2024-05-01Strong performance at a lower cost.
Key Capabilities
- MLA Architecture
- GPT-4 class
- Extremely cheap API
Innovations
- ★Multi-head Latent Attention (MLA)
- ★KV Cache Compression (93.3%)
- ★DeepSeekMoE Architecture
Command R+
2024-04-24104B Dense, RAG- und Tool-Use-optimiertes Open-Weight-Flagship (CC-BY-NC), 128k Kontext, citation-fähiges Retrieval. Starkes Grounding beim Zusammenfassen (Vectara HHEM 6.9% Halluzination für den 08-2024-Build), akademische Benchmarks dagegen schwach. Direkter Vorläufer der Command-A-Linie.
Key Capabilities
- 104B Dense
- RAG-optimiert
- Tool-Use
- Multilingual (10 Sprachen)
- 128k Context
Innovations
- ★Citation-fähiges RAG
- ★Open weights (CC-BY-NC)
Llama 3
2024-04-18Next generation open-source LLM with 8B and 70B variants.
Key Capabilities
- 8B/70B sizes
- Improved reasoning
- Enhanced coding
Innovations
- ★Improved Tokenizer
- ★GQA Across All Sizes
- ★Llama Guard Safety
Grok-1
2024-03-17314B MoE (86B active) open-weight model released under Apache 2.0. Largest open-weight model at the time of release.
Key Capabilities
- 314B total / 86B active MoE
- 8 experts, 2 active
- Apache 2.0 license
- Largest open-weight model at release
Innovations
- ★Largest Open-Weight MoE at Launch
- ★Apache 2.0 Full Weight Release
Claude 3 Opus
2024-03-04Haiku, Sonnet, and Opus models.
Key Capabilities
- Vision capabilities
- Near-human nuance
- Opus reasoning
Innovations
- ★Vision Capabilities
- ★Three-tier Model Family
- ★Near-human Reasoning
Mistral Large
2024-02-26Flagship model with top-tier reasoning capabilities.
Key Capabilities
- 32k Context
- Multi-lingual
- Strong reasoning
Innovations
- ★32k Context Window
- ★Multi-lingual Excellence
- ★Function Calling
Gemini 1.5 Pro
2024-02-15Mid-size multimodal model optimized for scaling.
Key Capabilities
- 1M+ Context window
- Video understanding
- In-context learning
Innovations
- ★1M+ Context Window
- ★MoE Architecture
- ★Video Understanding
DeepSeek-MoE
2024-01-01Mixture-of-Experts architecture.
Key Capabilities
- MoE efficiency
- High throughput
- Cost effective
Innovations
- ★MoE Efficiency
- ★Cost-effective Training
- ★High Throughput
Midjourney V6
2023-12-21Enhanced prompt accuracy and image coherence.
Key Capabilities
- Better prompt accuracy
- Improved coherence
- Advanced image prompting
Innovations
- ★Prompt Accuracy
- ★Image Coherence
- ★Advanced Prompting
Mixtral 8x7B
2023-12-11Sparse Mixture-of-Experts model.
Key Capabilities
- MoE Architecture
- GPT-3.5 level
- Open weights
Innovations
- ★Sparse MoE Architecture
- ★Router Network
- ★Expert Selection Mechanism
Gemini 1.0
2023-12-06Multimodal model built from the ground up.
Key Capabilities
- Native multimodal
- Nano/Pro/Ultra sizes
- AlphaCode 2 integration
Innovations
- ★Native Multimodality
- ★Ground-up Multimodal Training
- ★AlphaCode 2 Integration
GPT-4 Turbo
2023-11-06128k context window with knowledge up to April 2023, announced at DevDay.
Key Capabilities
- 128k Context
- JSON mode
- Cheaper pricing
Innovations
- ★128k Context Window
- ★JSON Mode
- ★Reproducible Outputs
DeepSeek-Coder
2023-11-02Open-source code generation model.
Key Capabilities
- Code specialization
- Open source
- Various sizes
Innovations
- ★Code Specialization
- ★Multiple Model Sizes
- ★Open Source Focus
Mistral 7B
2023-09-27Powerful open-weight model.
Key Capabilities
- Open weights
- Efficient attention
- High performance/size
Innovations
- ★Grouped Query Attention
- ★Sliding Window Attention
- ★Efficient Open Weights
Code Llama
2023-08-24Code-specialized Llama model for code generation and understanding.
Key Capabilities
- Code generation
- Infilling
- Instruction following
Innovations
- ★Code Specialization
- ★Infilling Capability
- ★Long Context Fine-tuning
Stable Diffusion XL
2023-07-26Major upgrade with native 1024x1024 resolution and improved generation.
Key Capabilities
- 1024x1024 native
- Two-stage pipeline
- Improved aesthetics
Innovations
- ★Higher Resolution
- ★Refiner Model
- ★Enhanced Prompt Understanding
Llama 2
2023-07-18Open-source LLM available for commercial use.
Key Capabilities
- 7B/13B/70B sizes
- Commercial license
- Chat fine-tuning
Innovations
- ★Open Commercial License
- ★RLHF Chat Models
- ★Safety Red-teaming
Midjourney V5
2023-03-15Major leap in photorealism and prompt understanding for AI art generation.
Key Capabilities
- Photorealistic images
- Improved prompts
- Higher resolution
Innovations
- ★Enhanced Photorealism
- ★Natural Language Understanding
- ★Artistic Quality
GPT-4
2023-03-14Multimodal model exhibiting human-level performance on various professional and academic benchmarks.
Key Capabilities
- Multimodal (Image inputs)
- Advanced reasoning
- 32k context window
Innovations
- ★Native Multimodal Input
- ★System Card Safety
- ★Advanced Vision Understanding
LLaMA
2023-02-24First open-weights large language model from Meta, released to researchers.
Key Capabilities
- Open weights
- 7B to 65B sizes
- Research access
Innovations
- ★Open-weight LLM
- ★Efficient Training
- ★Research Democratization
Stable Diffusion
2022-08-22Revolutionary open-source text-to-image model that democratized AI art.
Key Capabilities
- Text-to-Image
- Open source
- Local inference
- Community ecosystem
Innovations
- ★Latent Diffusion Model
- ★Open Source Democratization
- ★Diffusers Ecosystem