Gemini 2.0 Flash costs $0.100 per 1M input tokens and $0.400 per 1M output tokens. It features a 128k context window, a blended 3:1 production rate of $0.175/1M, and is optimized for real-time multimodal voice/video streaming, low-latency agent loops.
Gemini 2.0 Flash Token Pricing
Google Fast / MultimodalNext-gen multimodal workhorse with sub-second latency and 1M token context window.
Verified Specifications & Benchmark Data
| Developer / Provider | |
| Model Family | Gemini 2.0 |
| Blended 3:1 Rate (Production Benchmark) | $0.175 / 1M tokens |
| Batch API Discount (24hr SLA) | 50% off standard rate |
| Multimodal Vision | ✅ Supported (Images, Diagrams, OCR) |
| Function Calling / Structured Outputs | ✅ Native Tool Calling |
| MMLU Benchmark Score | 84.6% |
| Latency & Throughput Tier | Ultra-Fast (~110 t/s) |
| Optimal Architecture & Use Cases | Real-time multimodal voice/video streaming, low-latency agent loops |
Frequently Asked Questions about Gemini 2.0 Flash
How much does Gemini 2.0 Flash cost per 1M tokens?
Gemini 2.0 Flash pricing is set at $0.100 per 1 million input tokens and $0.400 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.
How much money does Gemini 2.0 Flash prompt caching save?
With prompt caching enabled, cached input tokens are discounted to $0.025/1M, saving 75% on repeated system prompts and document vectors.
What is the context window limit of Gemini 2.0 Flash?
Gemini 2.0 Flash has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.
Hosting & Cloud GPU Options for Gemini 2.0 Flash
Calculate whether serverless API or dedicated GPU hosting (RunPod / Together AI / Vultr) is more cost-effective for your volume.
Direct Matchups with Gemini 2.0 Flash
See how Gemini 2.0 Flash compares against other leading frontier and open-weights models in cost and latency.
Compare Google vs OpenAI pricing, context limits, and cost per 1M tokens.
Compare Google vs Groq (Meta) pricing, context limits, and cost per 1M tokens.
Compare Google vs Google pricing, context limits, and cost per 1M tokens.
Compare Google vs OpenAI pricing, context limits, and cost per 1M tokens.