Skip to content

Free tool

AI Model Comparison (2026)

Every major 2026 LLM side by side: context window, API price per 1M tokens, modality, and whether the weights are open. Filter by provider or open-weights, sort by price or context. Verified and cross-checked; long-context surcharges disclosed.

Want the editorial read on what is cheapest to run and how far open weights have closed the gap? See the Frontier Index →

All models. 29 of 29 models listed.. Cheapest input here is $0.1 per million, largest context 1050K.

Swipe for price & modality →

ModelContextIn $/1MOut $/1MWeightsModality
Qwen 3.7 MaxAlibaba · 2026-05The previous Max tier; strong tool-use + multilingual (esp. Chinese) reasoning1M$2.50*$7.50ProprietaryMultimodal
Qwen 3.8 MaxAlibaba · 2026-09-02Alibaba's current Max tier, 20% cheaper on both halves than Qwen 3.7 Max; thinking and non-thinking modes on one model ID1M$2*$6ProprietaryMultimodal
Claude Fable 5Anthropic · 2026-06-09Anthropic's most capable widely-released model; top-end reasoning + agentic work1M$10$50ProprietaryText + image in
Claude Fable 5.1Anthropic · 2026-09-01Anthropic's most capable model; demanding reasoning and long-horizon agentic work1M$10*$50ProprietaryText + image in
Claude Haiku 4.5Anthropic · 2025-10Fastest Claude with near-frontier intelligence at low cost200K$1$5ProprietaryText + image in
Claude Opus 4.8Anthropic · 2026Most capable Opus-tier model for complex reasoning + long-horizon agentic coding1M$5$25ProprietaryText + image in
Claude Opus 5Anthropic · 2026-07-24Complex agentic coding and enterprise work; supersedes Opus 4.81M$5*$25ProprietaryText + image in
Claude Sonnet 4.6Anthropic · 2026Best balance of speed and intelligence; strong agentic + coding workhorse1M$3*$15ProprietaryText + image in
Claude Sonnet 5Anthropic · 2026-06-30Best balance of speed and intelligence in the Claude 5 family1M$2*$10ProprietaryText + image in
DeepSeek V4DeepSeek · 2026-04-24Superseded by DeepSeek V4.1 Flash on 2026-09-10, and being phased out: DeepSeek routes deepseek-v4-pro requests to Flash from 2026-09-14 until a V4.1 Pro launches. Open-weights MoE (~32–37B active) with sparse attention1M$1.32*$3.96OpenText
DeepSeek V4.1 FlashDeepSeek · 2026-09-10552B MoE that activates only 8B parameters on input and 16B on output, with native visual understanding; DeepSeek's own benchmarks put it ahead of V4 Pro on quality, cost and speed at a quarter of the price1M$0.30*$1.20OpenMultimodal
Gemini 3.1 ProGoogle · 2026-02-19Frontier long-context multimodal reasoning; largest practical context in class1M$2*$12ProprietaryMultimodal (text/image/audio/video)
Gemini 3.5 FlashGoogle · 2026-05-20Google's legacy Flash tier: baseline speed for routine high-throughput work1M$1.50*$9ProprietaryMultimodal (text/image/audio/video)
Gemini 3.8 FlashGoogle · 2026Google's current Flash generation; even at its 2027 standard price it undercuts the Gemini 3.5 Flash it supersedes on output ($7.50 against $9.00)Not published$0.75*$3.75ProprietaryMultimodal (text/image/audio/video)
Llama 4 MaverickMeta · 2025-04-05Open-weights MoE; cheapest frontier-ish option, huge ecosystem + fine-tunability1M$0.20*$0.80OpenMultimodal (text + image)
Llama 4 ScoutMeta · 2025-04-05Smaller open-weights MoE (17B active, 16 experts) that fits on a single H100; the cheapest row in this ledger320K$0.10*$0.30OpenMultimodal (text + image)
Mistral Large 3Mistral AI · 2025-12-02Apache-2.0 open weights, 675B total / 41B active MoE; the cheapest way to self-host a Mistral of this size256K$0.50*$1.50OpenText + image
Mistral Medium 3.5Mistral AI · 2026-04-28Mistral's frontier-class model for agentic and coding work; open weights under a Modified MIT licence256K$1.50*$7.50OpenMultimodal
Kimi K3Moonshot AI · 2026Open-weights 2.8T-param MoE (104B active) under the Kimi K3 License; 1M context for long-horizon coding1M$3*$15OpenText (agentic)
GPT-5.4 miniOpenAI · 2026Cost-efficient mid-tier for high-volume tasks with solid reasoning400K$0.75$4.50ProprietaryText + image in
GPT-5.4 nanoOpenAI · 2026Cheapest OpenAI tier for simple, latency-sensitive, high-volume calls400K$0.20$1.25ProprietaryText + image in
GPT-5.5OpenAI · 2026-04-23Superseded as OpenAI's flagship by GPT-6 Astra on 2026-09-05; first OpenAI model with a 1M-token API context window1.05M$5*$30ProprietaryText + image in
GPT-5.6 LunaOpenAI · 2026-07-09Fastest, lowest-cost GPT-5.6 tier for high-volume work1.05M$0.20*$1.20ProprietaryText + image in
GPT-5.6 SolOpenAI · 2026-07-09Top GPT-5.6 tier for complex coding, research, and agentic work; below GPT-6 Astra since 2026-09-051.05M$4*$20ProprietaryText + image in
GPT-5.6 TerraOpenAI · 2026-07-09Balanced GPT-5.6 tier: near-GPT-5.5 quality at well under half the price1.05M$2*$12ProprietaryText + image in
GPT-6 AstraOpenAI · 2026-09-05OpenAI's most capable model: complex reasoning, coding, computer use, research, and document creation1.05M$10*$50ProprietaryText + image in
Grok 4.3xAI · 2026xAI's widest context at 1M tokens and its cheapest reasoning tier; native real-time X / web search1M$1.25*$2.50ProprietaryText + image in
Grok 4.6xAI · 2026-08-12xAI's frontier model for coding, agentic tasks and knowledge work; reasoning effort up to xhigh500K$2*$6ProprietaryText + image in
GLM-5Zhipu / Z.ai · 2026-02-11Open MIT-licensed frontier model; very strong coding + agentic value200K$1*$3.20OpenMultimodal

Standard pay-as-you-go list prices, USD per 1M tokens, as of 3 Sep 2026, for prompts up to ~200K tokens. * = tiered/long-context or promo pricing applies (hover the input price). Open-weights API prices are representative third-party hosted rates; self-hosting is free. This space moves weekly, so confirm with each provider before relying on a figure.

Frequently asked

Which 2026 model has the cheapest API pricing?

Among frontier-class options, third-party-hosted open-weights models are cheapest: Llama 4 Maverick runs about $0.2 / $0.8 per 1M tokens (input / output). Among major proprietary APIs, OpenAI's GPT-5.4 nano ($0.2 / $1.25) and Google's smaller Flash tiers are the lowest. DeepSeek V4 is unusually cheap for its capability, often discounted well below its $1.32 / $3.96 list rate.

Which model has the largest context window?

By the verified API context windows in this table, GPT-5.5 leads at ~1.05M tokens, just ahead of a ~1M cluster: Claude Opus 4.8 / Claude Sonnet 4.6 / Claude Fable 5, Gemini 3.1 Pro, Grok 4.3, and the open-weights DeepSeek V4 and Llama 4 Maverick. Google and Meta have publicly discussed larger windows (a Gemini tier up to ~2M, a Llama 4 Scout variant up to 10M), but those aren't reflected in these standard API rows.

What are the best open-weights models in mid-2026?

The strongest open-weights options are DeepSeek V4 (Apache-2.0), Meta's Llama 4 Maverick, Moonshot's Kimi K3, Zhipu's MIT-licensed GLM-5, and Mistral Large 3 (Apache-2.0). Several rival proprietary frontier models on coding and agentic tasks, and all can be self-hosted or run via many providers. Note that open weights and an open licence are not the same thing: DeepSeek V4, GLM-5 and Mistral Large 3 ship under OSI licences, while Kimi K3's weights are public under Moonshot's own Kimi K3 License.

Why are some prices shown with an asterisk and a footnote?

Several flagships use context-tiered pricing: the listed figure is the rate for prompts up to ~200K tokens, and the footnote shows the higher long-context rate. For example, Gemini 3.1 Pro lists $2 / $12 at the base tier but >200K billed $4/$18 (long-context tier); GPT-5.5 charges more on long prompts (>272K input billed 2x input / 1.5x output). Claude Opus 4.8, by contrast, has no long-context premium.

Is Grok 5 available yet?

Not as a generally available API as of 2026-09-03. Grok 5 has been widely discussed (rumored ~6T params, 1.5M context) but xAI had not published an official release or pricing at the time of writing, so this table uses Grok 4.3, the current GA flagship.

How to compare AI models

Comparing models is not about a single leaderboard number. The useful comparison lines up the things that actually affect your use: price per token, context-window size, speed, modality, and whether the weights are open.

Price is two numbers, not one

Input and output tokens are usually priced differently, so a model that looks cheap on input can be expensive if your workload generates long outputs. Compare a blended rate that reflects your own input-to-output mix, not just the headline input price.

Context and modality matter

A larger context window lets a model handle longer documents in one pass; multimodal support lets it read images or audio. A cheaper model with the right context or modality often beats a pricier one that lacks it.

Open versus closed

Open-weight models can be self-hosted for data control and zero per-token cost; closed API models are simplest and often the most capable. All figures here come from the hand-verified Frontier Ledger, so they stay consistent across the site.

Where this answer comes from

Last checked on Sep 3, 2026, and reviewed on a 30-day schedule. Figures come from the company that publishes them, never from another site’s summary. How we check this.

Related reading

Go deeper on how this works and what to pick.

Newsletter

Liked the tool? Get the signal.

One weekly email on the AI + hardware that actually matters, from the people who build these calculators.

Free · unsubscribe anytime · no spam.