Haiku 5.5 · over 100K
- Input: $0.50 / MTok
- Output: $2.50 / MTok
- Cache read: $0.05 / MTok
- 5-minute cache write: $0.625 / MTok
MODEL COMPARISON · VERIFIED OCTOBER 8, 2026
Two low-cost models with the same short-context token rates—but different long-context thresholds, reasoning controls and reported benchmark results.
Both models list $0.10 input and $0.50 output per million tokens, plus $0.01 cached input and $0.125 cache writes. Haiku 5.5 has stronger results on the selected benchmarks Anthropic published; GPT-6 Luna offers a 1.05M context window and does not enter its higher-price tier until prompts exceed 272K input tokens.
Standard API prices in USD per one million tokens. Provider and processing-tier prices can differ.
| Token category | Haiku 5.5 ≤100K prompt | GPT-6 Luna ≤272K input |
|---|---|---|
| Input | $0.10 | $0.10 |
| Output | $0.50 | $0.50 |
| Cached input / cache read | $0.01 | $0.01 |
| Cache write | $0.125 · 5 min | $0.125 |
| Batch processing | 50% off input/output | 50% of Standard rates |
Equal rate cards do not guarantee equal bills. Tokenizers, reasoning tokens, output length, retries and tool calls can change the cost per completed task.
Each provider applies its higher tier to the full qualifying request.
Between 100K and 272K prompt tokens, Haiku 5.5 is already in its higher tier while GPT-6 Luna remains at its short-context rate. This is a rate-card comparison only; task quality and token use still matter.
One standard request, no cache hits or tool fees.
Haiku 5.5: 200K × $0.50/MTok + 50K × $2.50/MTok.
GPT-6 Luna: 200K × $0.10/MTok + 50K × $0.50/MTok.
Haiku’s list cost in this specific token-count example. This is not a quality-adjusted comparison.
These figures come from Anthropic’s Haiku 5.5 announcement. Vendor-reported benchmark comparisons should be validated on your own tasks.
| Benchmark | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| GDPval-AA v2.1 | 1620 | 1437 |
| AA-Briefcase v1.1 | 1578 | 1336 |
| OSWorld 2.1 · offline subset | 72.4% | 48.9% |
| Terminal-Bench 4.0 | 39.2% | 16.4% |
| FrontierCode 1.1 · main | 46.4% | 42.4% |
| Chartography · no tools | 46.4% | 29.1% |
Both target high-volume, cost-sensitive work.
best model = lowest total cost per accepted resultMeasure success rate, p50/p95 latency, input and output tokens, retries, cache hit rate and any paid tool calls on the same evaluation set.
Not in their short-context tiers: both list $0.10 input, $0.50 output, $0.01 cached input and $0.125 cache writes per million tokens. Their higher tiers begin at different thresholds.
OpenAI documents a 1,050,000-token context window for GPT-6 Luna. Haiku 5.5’s 100K figure on this page is a pricing threshold, not necessarily its maximum context limit.
OpenAI says prompts with more than 272K input tokens use 2× input and cache rates and 1.5× output pricing for the full request.
The comparison table reproduces results published by Anthropic. Treat them as vendor-reported data and run a representative evaluation before selecting a model.
Anthropic lists claude-haiku-5-5; OpenAI lists gpt-6-luna.