Batch API pricing & savings
Batch APIs run requests asynchronously (results usually within 24 hours) in exchange for a lower price. Good fits: evaluations, classification, summarizing archives, embeddings backfills. Bad fits: anything a user is waiting on.
Estimate your monthly bill with this discount →
| Model | Normal input /1M | Batch input /1M | Saving |
|---|---|---|---|
| Claude Fable 5 | $10.00 | $5.00 | 50% |
| Claude Fable 5.1 | $10.00 | $5.00 | 50% |
| Claude Haiku 4.5 | $1.00 | $0.50 | 50% |
| Claude Mythos 5 | $10.00 | $5.00 | 50% |
| Claude Mythos 5.1 | $10.00 | $5.00 | 50% |
| Claude Opus 4.5 | $5.00 | $2.50 | 50% |
| Claude Opus 4.6 | $5.00 | $2.50 | 50% |
| Claude Opus 4.7 | $5.00 | $2.50 | 50% |
| Claude Opus 4.8 | $5.00 | $2.50 | 50% |
| Claude Opus 5 | $5.00 | $2.50 | 50% |
| Claude Opus 5.5 | $4.00 | $2.00 | 50% |
| Claude Sonnet 4.5 | $3.00 | $1.50 | 50% |
| Claude Sonnet 4.6 | $3.00 | $1.50 | 50% |
| Claude Sonnet 5 | $2.00 | $1.00 | 50% |
| Claude Sonnet 5.5 | $2.00 | $1.00 | 50% |
| Gemini 2.5 Flash | $0.30 | $0.15 | 50% |
| Gemini 2.5 Flash Lite | $0.10 | $0.05 | 50% |
| Gemini 2.5 Pro | $1.25 | $0.625 | 50% |
| Gemini 3 Flash Preview | $0.50 | $0.25 | 50% |
| Gemini 3.1 Flash Lite | $0.25 | $0.125 | 50% |
| Gemini 3.1 Pro Preview | $2.00 | $1.00 | 50% |
| Gemini 3.5 Flash | $1.50 | $0.75 | 50% |
| Gemini 3.5 Flash Lite | $0.30 | $0.15 | 50% |
| Gemini 3.6 Flash | $0.75 | $0.375 | 50% |
| Gemini 3.7 Flash | $0.75 | $0.375 | 50% |
| Gemini 3.8 Flash | $0.75 | $0.375 | 50% |
| Gemini Flash (latest) | $0.75 | $0.375 | 50% |
| Gemini Flash Lite (latest) | $0.30 | $0.15 | 50% |
| Gemini Pro (latest) | $2.00 | $1.00 | 50% |
| GPT-4.1 | $2.00 | $1.00 | 50% |
| GPT-4.1 mini | $0.40 | $0.20 | 50% |
| GPT-4o | $2.50 | $1.25 | 50% |
| GPT-4o mini | $0.15 | $0.075 | 50% |
| GPT-5 | $1.25 | $0.625 | 50% |
| GPT-5 mini | $0.25 | $0.125 | 50% |
| GPT-5 nano | $0.05 | $0.025 | 50% |
| GPT-5.1 | $1.25 | $0.625 | 50% |
| GPT-5.2 | $1.75 | $0.875 | 50% |
| GPT-5.4 | $2.50 | $1.25 | 50% |
| GPT-5.4 mini | $0.75 | $0.375 | 50% |
| GPT-5.4 nano | $0.20 | $0.10 | 50% |
| GPT-5.5 | $5.00 | $2.50 | 50% |
| GPT-5.6 Luna | $0.20 | $0.10 | 50% |
| GPT-5.6 Sol | $4.00 | $2.00 | 50% |
| GPT-5.6 Terra | $2.00 | $1.00 | 50% |
| GPT-6 Astra | $10.00 | $5.00 | 50% |
| GPT-6 Luna | $0.10 | $0.05 | 50% |
| GPT-6 Sol | $2.00 | $1.00 | 50% |
| GPT-6.1 Sol | $2.00 | $1.00 | 50% |
| o3 | $2.00 | $1.00 | 50% |
| Grok 4.20 Non Reasoning | $1.25 | $1.00 | 20% |
| Grok 4.20 Reasoning | $1.25 | $1.00 | 20% |
| Grok 4.3 | $1.25 | $1.00 | 20% |
FAQ
Is batch output also discounted?
Usually yes, by the same share as input. Each model page lists both batch input and batch output prices.
Can I combine batch with prompt caching?
Some providers allow it and bill cached tokens at a reduced rate inside batches; check the provider's documentation for current rules.
Which providers have no batch price here?
Models without a published batch price in our data are left out of this table.