At or below 100K
- Input: $0.10 / MTok
- Output: $0.50 / MTok
- Cache read: $0.01 / MTok
- Best unit economics for short, high-volume work
API PRICING GUIDE · VERIFIED OCTOBER 8, 2026
The complete cost guide for both prompt-length tiers, cache reads and writes, Batch API and US-only inference—with worked examples you can audit.
For prompts over 100,000 tokens, Haiku 5.5 rises to $0.50 per million input tokens and $2.50 per million output tokens. The threshold is evaluated per request, not from your monthly token total.
USD per one million tokens on the Claude Platform.
| Token category | Prompt ≤100K | Prompt >100K |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache read / hit | $0.01 | $0.05 |
| Cache write · 5 minutes | $0.125 | $0.625 |
| Cache write · 1 hour | $0.20 | $1.00 |
The short-prompt tier is one fifth of the long-prompt tier for each token category. Cache hits are one tenth of the tier’s base input rate.
Open the interactive calculatorIt is not a monthly volume threshold.
Anthropic describes the lower tier as prompts “up to 100,000 tokens” and the higher tier as prompts “over 100,000 tokens.” Exactly 100,000 prompt tokens therefore remain in the lower tier.
Standard global pricing, before any cache or Batch adjustments.
10K input + 1K output in the short tier: 10,000 × $0.10/MTok + 1,000 × $0.50/MTok.
100K input + 20K output in the short tier. At 1,000 requests, that is $20.
200K input + 50K output in the long tier. At 1,000 requests, that is $225.
request cost = (input ÷ 1M × input rate) + (output ÷ 1M × output rate) + cache costsMultiply the request result by request volume. Add any paid server-side tool usage separately.
Pricing modifiers may stack, so model the full request configuration.
Batch API discount on input and output tokens for asynchronous processing.
Haiku cache-hit multiplier versus the applicable base input rate.
US-only inference multiplier for Claude 4.6 and later models; global routing uses standard pricing.
A five-minute cache write costs 1.25× the base input rate and a one-hour write costs 2×. Caching pays off when reused prompt content avoids repeated full-price input processing.
Prompts up to 100K cost $0.10 per million input tokens and $0.50 per million output tokens. Prompts over 100K cost $0.50 input and $2.50 output.
Yes. Anthropic’s rate card says “up to 100,000 tokens” for the lower tier and “over 100,000 tokens” for the higher tier.
In the short tier, a five-minute write is $0.125/MTok, a one-hour write is $0.20/MTok and a cache hit is $0.01/MTok. In the long tier they are $0.625, $1.00 and $0.05 respectively.
Yes. Anthropic documents a 50% input and output token discount for asynchronous Batch API workloads.
Do not assume so. Partner-operated platforms can have their own regional pricing. Check the provider’s current rate card before budgeting.
claude-haiku-5-5 on the Claude Platform.