GPT-4o Mini by OpenAI costs $0.150 per 1 Million input tokens and $0.600 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $0.263/1M tokens.
GPT-4o Mini
Ultra-affordable small model replacing GPT-3.5 Turbo with significantly superior reasoning and vision.
Verified Specifications & Benchmark Data
| Developer / Provider | OpenAI |
| Model Family | GPT-4 |
| Blended 3:1 Rate (Production Benchmark) | $0.263 / 1M tokens |
| Batch API Discount (24hr SLA) | 50% off standard rate |
| Multimodal Vision | ✅ Supported (Images, Diagrams, OCR) |
| Function Calling / Structured Outputs | ✅ Native Tool Calling |
| MMLU Benchmark Score | 82.0% |
| Latency & Throughput Tier | Ultra-Fast (~85 t/s) |
| Optimal Architecture & Use Cases | High-volume classification, customer support bots, lightweight summarization |
Frequently Asked Questions about GPT-4o Mini
How much does GPT-4o Mini cost per 1M tokens?
GPT-4o Mini pricing is set at $0.150 per 1 million input tokens and $0.600 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.
How much money does GPT-4o Mini prompt caching save?
With prompt caching enabled, cached input tokens are discounted to $0.075/1M, saving 50% on repeated system prompts and document vectors.
What is the context window limit of GPT-4o Mini?
GPT-4o Mini has a maximum context window of 128k (128,000 tokens), supporting up to 16,384 completion tokens per response.
Direct Matchups with GPT-4o Mini
See how GPT-4o Mini compares against other leading frontier and open-weights models in cost and latency.
Compare OpenAI vs OpenAI pricing, context limits, and cost per 1M tokens.
Compare OpenAI vs Groq (Meta) pricing, context limits, and cost per 1M tokens.
Compare OpenAI vs Google pricing, context limits, and cost per 1M tokens.
Compare OpenAI vs Google pricing, context limits, and cost per 1M tokens.