Home  /  Models  /  Gemini 1.5 Flash
⚡ Quick Pricing Summary

Gemini 1.5 Flash by Google costs $0.075 per 1 Million input tokens and $0.300 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $0.131/1M tokens.

Google Fast / Lightweight

Gemini 1.5 Flash

Lightweight model with 1M context window engineered for high speed and cost efficiency.

Calculate Spend in App →
Input Token Price
$0.075
Per 1,000,000 tokens
Output Token Price
$0.300
Per 1,000,000 tokens
Prompt Caching Rate
$0.019
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider Google
Model Family Gemini 1.5
Blended 3:1 Rate (Production Benchmark) $0.131 / 1M tokens
Batch API Discount (24hr SLA) 50% off standard rate
Multimodal Vision ✅ Supported (Images, Diagrams, OCR)
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 78.9%
Latency & Throughput Tier Ultra-Fast (~100 t/s)
Optimal Architecture & Use Cases High-frequency API calls, audio processing, bulk document summarization

Frequently Asked Questions about Gemini 1.5 Flash

How much does Gemini 1.5 Flash cost per 1M tokens?

Gemini 1.5 Flash pricing is set at $0.075 per 1 million input tokens and $0.300 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.

How much money does Gemini 1.5 Flash prompt caching save?

With prompt caching enabled, cached input tokens are discounted to $0.019/1M, saving 75% on repeated system prompts and document vectors.

What is the context window limit of Gemini 1.5 Flash?

Gemini 1.5 Flash has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Direct Matchups with Gemini 1.5 Flash

See how Gemini 1.5 Flash compares against other leading frontier and open-weights models in cost and latency.