Claude 3.5 Haiku by Anthropic costs $0.800 per 1 Million input tokens and $4.00 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $1.600/1M tokens.
Claude 3.5 Haiku
High-speed intelligence matching previous-generation flagship performance at a fraction of latency.
Verified Specifications & Benchmark Data
| Developer / Provider | Anthropic |
| Model Family | Claude 3.5 |
| Blended 3:1 Rate (Production Benchmark) | $1.600 / 1M tokens |
| Batch API Discount (24hr SLA) | 50% off standard rate |
| Multimodal Vision | ❌ Text Only |
| Function Calling / Structured Outputs | ✅ Native Tool Calling |
| MMLU Benchmark Score | 75.2% |
| Latency & Throughput Tier | Ultra-Fast (~95 t/s) |
| Optimal Architecture & Use Cases | Real-time agents, code completion, fast document scanning |
Frequently Asked Questions about Claude 3.5 Haiku
How much does Claude 3.5 Haiku cost per 1M tokens?
Claude 3.5 Haiku pricing is set at $0.800 per 1 million input tokens and $4.00 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.
How much money does Claude 3.5 Haiku prompt caching save?
With prompt caching enabled, cached input tokens are discounted to $0.080/1M, saving 90% on repeated system prompts and document vectors.
What is the context window limit of Claude 3.5 Haiku?
Claude 3.5 Haiku has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.
Direct Matchups with Claude 3.5 Haiku
See how Claude 3.5 Haiku compares against other leading frontier and open-weights models in cost and latency.
Compare Anthropic vs Google pricing, context limits, and cost per 1M tokens.
Compare Anthropic vs OpenAI pricing, context limits, and cost per 1M tokens.
Compare Anthropic vs DeepSeek pricing, context limits, and cost per 1M tokens.
Compare Anthropic vs Groq (Meta) pricing, context limits, and cost per 1M tokens.