Home  /  Models  /  GPT-4o (Omni)
⚡ Quick Answer (TL;DR)

GPT-4o (Omni) costs $2.50 per 1M input tokens and $10.00 per 1M output tokens. It features a 128k context window, a blended 3:1 production rate of $4.375/1M, and is optimized for complex enterprise workflows, high-precision coding, multi-modal analysis.

GPT-4o (Omni) Token Pricing

OpenAI Flagship

OpenAI's flagship versatile multimodal model with native audio, vision, and fast reasoning capabilities.

Simulate GPT-4o (Omni) Bill →
Input Price
$2.50
per 1 Million tokens
Output Price
$10.00
per 1 Million tokens
Prompt Caching
$1.250
Saves up to 50%
Context Window
128k
Max output: 16k tokens

Verified Specifications & Benchmark Data

Developer / Provider OpenAI
Model Family GPT-4
Blended 3:1 Rate (Production Benchmark) $4.375 / 1M tokens
Batch API Discount (24hr SLA) 50% off standard rate
Multimodal Vision ✅ Supported (Images, Diagrams, OCR)
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 88.7%
Latency & Throughput Tier Fast (~45 t/s)
Optimal Architecture & Use Cases Complex enterprise workflows, high-precision coding, multi-modal analysis

Frequently Asked Questions about GPT-4o (Omni)

How much does GPT-4o (Omni) cost per 1M tokens?

GPT-4o (Omni) pricing is set at $2.50 per 1 million input tokens and $10.00 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.

How much money does GPT-4o (Omni) prompt caching save?

With prompt caching enabled, cached input tokens are discounted to $1.250/1M, saving 50% on repeated system prompts and document vectors.

What is the context window limit of GPT-4o (Omni)?

GPT-4o (Omni) has a maximum context window of 128k (128,000 tokens), supporting up to 16,384 completion tokens per response.

Hosting & Cloud GPU Options for GPT-4o (Omni)

Calculate whether serverless API or dedicated GPU hosting (RunPod / Together AI / Vultr) is more cost-effective for your volume.

Request Custom Audit →

Direct Matchups with GPT-4o (Omni)

See how GPT-4o (Omni) compares against other leading frontier and open-weights models in cost and latency.