LLM API Pricing Comparison 2026

Compare pricing across 64 models from 14 providers including GPT-5.6 Sol, Claude Fable 5, Gemini 3 Pro, Kimi K3, DeepSeek V4, and Grok. All prices per 1 million tokens, updated daily.

Get notified when prices change
64 models
Provider
Mistral NemoCheapest tier. 13B, built with NVIDIA, Apache 2.0.Mistral$0.020$0.030
Amazon Nova MicroCheapest tier. Text-only optimized for classification/routing.Amazon$0.035$0.14
GPT-5 nanoLowest-cost model in the GPT-5 line. Classification/extraction.OpenAI$0.050$0.40
Llama 3.1 8B Instant (Groq)Cheapest hosted model. Lowest verified output price in market ($0.08/M).Groq$0.050$0.080
Qwen TurboCheapest text tier. 50% batch discount.Alibaba Qwen$0.050$0.20
Amazon Nova LiteMid-tier balanced multimodal. Supports 300K context.Amazon$0.060$0.24
GPT-OSS 120B (Groq)OpenAI's open-weight model on Groq. Headline low-latency example.Groq$0.075$0.30
Llama 4 Scout (Groq)Meta 109B MoE. ~3.1M context on Groq. Ultra-low latency LPU.Groq$0.11$0.34
Mistral Small 4Hybrid reasoning, MoE 119B total / ~6.5B active. Apache 2.0. Day-0 NVIDIA NIM.Mistral$0.15$0.60
Command R 08-2024Budget tier. Cheapest production model.Cohere$0.15$0.60
MiniMax-M2.5Discount text tier. 197K context. Cheapest MiniMax text model.MiniMax$0.15$0.90
Grok 4 Fast2M context (largest in xAI lineup). Batch pricing; live/sync $0.40/$1.00.xAI$0.20$0.50
Grok Code Fast 1Coding specialist. 70.8% SWE-Bench Verified. For agentic coding IDEs.xAI$0.20$1.50
Qwen3 235B (Together)Flagship Qwen3 MoE. Thinking-mode capable.Together AI$0.20$0.60
Qwen3 VL PlusVision. Tiered >32K tokens. Thinking + non-thinking.Alibaba Qwen$0.20$1.60
Qwen3 235B (Fireworks)Qwen3 flagship MoE. Recently reduced from $0.60/$1.60.Fireworks AI$0.22$0.65
GPT-5.2 miniSmall/cheap tier of the 5.2 family; same modalities.OpenAI$0.25$2.00
DeepSeek V4-FlashCheaper MoE variant. 284B total / ~13B active. Same peak/off-peak schedule as V4-Pro.DeepSeek$0.27$1.10
Qwen3 32B (Groq)Mid-size Qwen3 dense model.Groq$0.29$0.39
Gemini 2.5 FlashCheap/fast prior gen. Still sold, flagged for deprecation.Google$0.30$2.50
Codestral 25.08Coding specialist. 82+ languages, fill-in-the-middle.Mistral$0.30$0.90
MiniMax M3 (Together)Chinese frontier MoE hosted on Together. 428B total / ~23B active.Together AI$0.30$0.30
Qwen3 Coder FlashCost-tier coder. Tiered >32K tokens.Alibaba Qwen$0.30$1.50
MiniMax-M3Current flagship. 428B total / ~23B active MoE. Open weights on HuggingFace since 2026-06-07. 1M context, MiniMax Sparse Attention.MiniMax$0.30$1.20
MiniMax-M2.7Mid-generation text model. 205K context.MiniMax$0.30$1.20
Devstral 2Coding/agentic dev model.Mistral$0.40$2.00
Qwen PlusBalanced workhorse. 1M context. Thinking-mode output $4/1M.Alibaba Qwen$0.40$1.20
Gemini 3 FlashBalanced speed/price. Same multimodal set as 3 Pro.Google$0.50$3.00
Magistral SmallMagistral Small 1.2. Image-capable reasoning model.Mistral$0.50$1.50
Llama 4 Maverick (Groq)Flagship Llama 4 hosted day-zero on Groq.Groq$0.50$0.65
Llama 3.3 70B Versatile (Groq)Versatile default. Production-tier on Groq.Groq$0.59$0.79
DeepSeek V3.1 (Together)Flagship DeepSeek MoE 671B. Together's most popular reasoning model.Together AI$0.60$1.70
Llama 4 Maverick (Together)Meta's flagship MoE. Natively multimodal, 1M context.Together AI$0.62$0.99
Amazon Nova ProFrontier multimodal tier with video support. ~4.3x cheaper than comparable frontier.Amazon$0.80$3.20
Llama 4 Maverick Basic (Fireworks)Meta Maverick 400B MoE. 1M context. 'basic' tier = lowest cost.Fireworks AI$0.85$1.20
Llama 3.3 70B (Fireworks)Still widely-used dense 70B workhorse.Fireworks AI$0.88$0.88
Kimi K2.7 CodeCoding-focused. Supports long thinking, tool calls, JSON mode.Moonshot AI$0.95$4.00
Kimi K2.6General-purpose. Thinking + non-thinking modes, agent tasks, web search.Moonshot AI$0.95$4.00
Claude Haiku 4.5Fast/cheap tier. First Haiku with reasoning (extended thinking).Anthropic$1.00$5.00
Qwen3 Coder PlusCoding flagship. Tiered >32K tokens. Prompt-caching discount.Alibaba Qwen$1.00$5.00
o4-miniCheaper/faster reasoning model. Workhorse for math + code.OpenAI$1.10$4.40
DeepSeek R1 (Fireworks)Reasoning model on V3 architecture. Full chain-of-thought.Fireworks AI$1.10$1.10
Qwen3 MaxStable production flagship. Tiered: $1.20/$6 (0-32K), $2.40/$12 (32-128K), $3/$15 (128-256K).Alibaba Qwen$1.20$6.00
GPT-5.2Prior flagship. Multimodal in/out, 400K context, configurable reasoning effort.OpenAI$1.25$10.00
GPT-5.1Prior flagship; still sold. Predecessor to 5.2.OpenAI$1.25$10.00
Gemini 2.5 ProPrior flagship, still on the API. Superseded by Gemini 3 Pro.Google$1.25$10.00
Gemini 3.5 FlashNewest Flash. Improved reasoning (thinking) over 3 Flash.Google$1.50$9.00
Mistral Medium 3.5Dense ~128B params (not MoE). Flagship of the Medium tier.Mistral$1.50$7.50
DeepSeek V4-ProOff-peak pricing. Peak hours (~09:00-18:00 Beijing) ~2x. Unified chat + reasoning MoE.DeepSeek$1.74$3.48
o3Reasoning model. Heavy thinking trace, still on the API.OpenAI$2.00$8.00
Gemini 3 ProTop-tier Gemini. Native multimodal incl. video; 2x price >200K tokens.Google$2.00$12.00
Grok 4.5Flagship. Tiered: >200K tokens costs $4 in / $12 out per 1M. ~2x token efficiency vs Grok 4.xAI$2.00$6.00
Mistral Large 2512Current alias target. Apache 2.0, MoE ~675B.Mistral$2.00$6.00
Magistral MediumNative reasoning model. Latest sub-version is Magistral Medium 1.2.Mistral$2.00$5.00
Command AFlagship. 111B params, tool-use / agentic focus.Cohere$2.50$10.00
Command R+ 08-2024Still sold as legacy alternative to Command A.Cohere$2.50$10.00
Amazon Nova PremierAmazon's most capable model. 1M context. Promotional $2.50/$12.50 through Aug 31 2026.Amazon$2.50$12.50
Qwen3.7 MaxTop-of-line proprietary flagship. 1M context, hybrid thinking. Limited-time 50%-off promo.Alibaba Qwen$2.50$7.50
Claude Sonnet 4.6Balanced workhorse. 1M context GA at standard pricing.Anthropic$3.00$15.00
Grok 4Still sold. Always-on reasoning (no off switch).xAI$3.00$15.00
Kimi K3Flagship. 2.8T params (largest open-weight model ever). Kimi Delta Attention architecture. Hybrid reasoning.Moonshot AI$3.00$15.00
GPT-5.6 SolCurrent flagship — part of the GPT-5.6 family (Sol / Luna / Terra). Sol is the workhorse reasoning variant for coding, science, cybersecurity, and computer use. ~54% more token-efficient for coding. 2x input / 1.5x output surcharge above 272K tokens.OpenAI$5.00$30.00
Claude Opus 4.8Top-tier agentic coding + enterprise model. 1M context default.Anthropic$5.00$25.00
Claude Fable 5Anthropic's most capable widely-released model for demanding reasoning and long-horizon agentic work. Sibling to Claude Mythos 5. 2x/1.5x surcharge applies above 200K input tokens. Globally re-available since Jul 1, 2026 after temporary U.S. export-control restriction.Anthropic$10.00$50.00

Frequently Asked Questions

What is the cheapest LLM API in 2026?

Mistral Nemo is the cheapest at $0.02/1M input tokens. For frontier-class open-weight models, MiniMax-M3 at $0.30/$1.20 per 1M tokens is exceptional value. Among major Western providers, Amazon Nova Micro ($0.035/$0.14) and Google Gemini 2.5 Flash ($0.30/$2.50) are the cheapest.

How much does GPT-5.6 Sol cost per token?

GPT-5.6 Sol costs $5.00 per 1 million input tokens and $30.00 per 1 million output tokens. The flagship of the GPT-5.6 family (Sol/Luna/Terra), released July 9, 2026. Prompts above 272K tokens are billed at 2x input / 1.5x output. Cached input is 90% cheaper at $0.50/1M tokens.

How much does Claude Fable 5 cost?

Claude Fable 5 costs $10.00/1M input and $50.00/1M output tokens with a 1M context window. Released June 9, 2026 alongside Claude Mythos 5. Prompt caching drops cached input to $1.00/1M. Note: Fable 5 was temporarily restricted under U.S. export controls mid-June 2026 and returned globally on July 1, 2026.

Which LLM has the largest context window?

Groq's hosted Llama 4 Scout leads with ~3.1M tokens (Meta's theoretical max is 10M). xAI Grok 4 Fast offers 2M tokens. OpenAI GPT-5.2 offers 400K, while Google Gemini 3 Pro, Anthropic Claude Opus 4.8, Moonshot Kimi K3, and Alibaba Qwen3.7 Max all offer 1M tokens.

How do I calculate my LLM API cost?

Multiply your average input tokens per request by the input price, add your output tokens times the output price, then multiply by your daily request count and 30 days. Use our calculator for instant comparisons across all providers.