LLM API pricing, side by side
Token prices for six providers, taken from their official pricing pages and nothing else. Standard tier, USD per 1 million tokens, with the caching and long-context caveats that quietly change your bill.
Runs entirely in your browser. Nothing you enter leaves this page.
Every number verified against the linked source on 2026-08-31. The spread is real: Mistral Small 4 costs $0.75 per combined 1M in + 1M out, GPT-5.5 Pro costs $210.00.
Anthropic
official pricing page| Model | Input /1M | Cached /1M | Output /1M | 1M in + 1M out | Context |
|---|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $1.00 | $50.00 | $60.00 | 1M |
| Claude Mythos 5 | $10.00 | $1.00 | $50.00 | $60.00 | 1M |
| Claude Opus 4.8 | $5.00 | $0.50 | $25.00 | $30.00 | 1M |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | $30.00 | 1M |
| Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 | $18.00 | 1M |
| Claude Sonnet 5 | $2.00 | $0.20 | $10.00 | $12.00 | 1M |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | $6.00 | - |
- Cached input is the cache read price (0.1x input). Cache writes cost extra: 1.25x input for 5-minute TTL, 2x for 1-hour TTL.
- Batch API: 50 percent off input and output.
- Claude Sonnet 5's $2 in / $10 out was announced as introductory pricing through 2026-08-31. Anthropic has since made it the standard price and cancelled the increase to $3 / $15 that was scheduled for 2026-09-01.
- Claude 4.6 and later models, plus Claude Mythos Preview, include the full 1M-token context window at standard pricing. Anthropic states the rule by generation rather than by model list.
- Pinning inference to the US with the inference_geo parameter costs 1.1x on every token category, including cache writes and reads, on Claude 4.6 and later. Global routing is the default and is priced as shown.
OpenAI
official pricing page| Model | Input /1M | Cached /1M | Output /1M | 1M in + 1M out | Context |
|---|---|---|---|---|---|
| GPT-5.5 Protiered | $30.00 | - | $180.00 | $210.00 | - |
| GPT-5.5tiered | $5.00 | $0.50 | $30.00 | $35.00 | - |
| GPT 5.6 Sol | $4.00 | $0.40 | $20.00 | $24.00 | - |
| GPT-5.4tiered | $2.50 | $0.25 | $15.00 | $17.50 | - |
| GPT 5.6 Terra | $2.00 | $0.20 | $12.00 | $14.00 | - |
| GPT-5.3-Codex | $1.75 | $0.175 | $14.00 | $15.75 | - |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 | $5.25 | - |
| GPT-5.4 nano | $0.20 | $0.02 | $1.25 | $1.45 | - |
| GPT 5.6 Luna | $0.20 | $0.02 | $1.20 | $1.40 | - |
- The flagship GPT-5.6 models each carry a separate long-context band on the pricing page, at 2x input and 1.5x output. OpenAI does not state the token threshold. Prior-generation rows now sit behind the page's "All models" view rather than on the flagship table.
- Batch and Flex: 50 percent off. Fast mode, renamed from Priority processing on 2026-07-30, costs 2x standard.
| Model | Input /1M | Cached /1M | Output /1M | 1M in + 1M out | Context |
|---|---|---|---|---|---|
| Gemini 3.1 Pro (preview)tiered | $2.00 | $0.20 | $12.00 | $14.00 | - |
| Gemini 3.5 Flash | $1.50 | $0.15 | $9.00 | $10.50 | - |
| Gemini 2.5 Protiered | $1.25 | $0.125 | $10.00 | $11.25 | - |
| Gemini 3.6 Flash | $0.75 | $0.075 | $3.75 | $4.50 | - |
| Gemini 3.7 Flash | $0.75 | $0.075 | $3.75 | $4.50 | - |
| Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 | $2.80 | - |
- Gemini 3.1 Pro and 2.5 Pro charge higher rates for prompts above 200K tokens (3.1 Pro: $4.00 in / $18.00 out).
- Flash models charge more for audio input than text input. Rates shown are text/image/video.
- Context caching storage costs $1.00 per 1M tokens per hour on top of cached-input rates, halved to $0.50 through 2026-12-31 on the promotionally-priced Flash models.
- Gemini 3.6 Flash and 3.7 Flash are promotionally priced at $0.75 in / $3.75 out through 2026-12-31. Both rise to $1.50 / $7.50 on 2027-01-01.
DeepSeek
official pricing page| Model | Input /1M | Cached /1M | Output /1M | 1M in + 1M out | Context |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | $0.66 | $0.022 | $1.98 | $2.64 | 1M |
| DeepSeek V4 Flash | $0.22 | $0.007 | $0.66 | $0.88 | 1M |
| Deepseek V4 Flash Vision Exp | $0.22 | $0.007 | $0.66 | $0.88 | - |
- Cached input is the cache-hit price; cache misses pay the normal input rate.
- Peak/off-peak pricing took effect 2026-08-16, and both halves rose at the same time. Rates shown are OFF-PEAK, which applies 17 of 24 hours; peak hours (01:00-04:00 and 06:00-10:00 UTC) cost exactly 2x.
- The deepseek-chat and deepseek-reasoner aliases were deprecated on 2026-07-24 and no longer appear on DeepSeek's pricing page.
Mistral
official pricing page| Model | Input /1M | Cached /1M | Output /1M | 1M in + 1M out | Context |
|---|---|---|---|---|---|
| Magistral Medium | $2.00 | - | $5.00 | $7.00 | - |
| Mistral Medium 3.5 | $1.50 | - | $7.50 | $9.00 | - |
| Mistral Large 3 | $0.50 | - | $1.50 | $2.00 | - |
| Codestral | $0.30 | - | $0.90 | $1.20 | - |
| Mistral Small 4 | $0.15 | - | $0.60 | $0.75 | - |
- Mistral Large 3 is priced below the newer Mistral Medium 3.5 on the official page ($0.50 / $1.50 against $1.50 / $7.50). Verified twice on 2026-06-10 and again on 2026-08-19, because the ordering reads like an error and is not one.
- Batch API: 50 percent off. Prompt caching cuts input cost by 90 percent, and regional inference adds 10 percent.
| Model | Input /1M | Cached /1M | Output /1M | 1M in + 1M out | Context |
|---|---|---|---|---|---|
| Grok 4.5tiered | $2.00 | $0.30 | $6.00 | $8.00 | 500K |
| Grok 4.6tiered | $2.00 | $0.50 | $6.00 | $8.00 | 500K |
| Grok 4.20 (reasoning)tiered | $1.25 | $0.20 | $2.50 | $3.75 | 1M |
| Grok 4.3tiered | $1.25 | $0.20 | $2.50 | $3.75 | 1M |
| Grok Build 0.1tiered | $1.00 | $0.20 | $2.00 | $3.00 | 256K |
- Batch API: 20 to 50 percent off standard rates, varies per model.
- Every Grok model has a long-context tier. xAI publishes one threshold for all of them, 200,000 tokens, above which input, cached input and output each cost exactly 2x the rate shown.
How to read this table honestly
- Cached input is not one thing. Anthropic and DeepSeek list a cache-read price; Anthropic also bills cache writes (1.25x input), Google bills cache storage per hour. Two providers with the same "cached" number can produce different bills.
- "Tiered" means the listed price is the floor. OpenAI's GPT-5.5/5.4 and Google's Pro models charge roughly double for long-context requests, which is exactly what agent workflows produce.
- Batch discounts are large. Anthropic, OpenAI, and Mistral all offer 50 percent off for async batch workloads; xAI 20 to 50 percent.
- Excluded providers are excluded for a reason. Meta Llama API, Cohere, and Amazon Nova render their prices client-side, so they could not be verified from source. No number on this page is quoted from a third-party blog.
