Home  /  Models  /  Llama 3.3 70B (Groq LPU)
⚡ Quick Pricing Summary

Llama 3.3 70B (Groq LPU) by Groq (Meta) costs $0.590 per 1 Million input tokens and $0.790 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $0.640/1M tokens.

Groq (Meta) Ultra-Fast Open Weights

Llama 3.3 70B (Groq LPU)

Meta's flagship 70B open model hosted on Groq LPUs delivering 280+ tokens per second.

Calculate Spend in App →
Input Token Price
$0.590
Per 1,000,000 tokens
Output Token Price
$0.790
Per 1,000,000 tokens
Prompt Caching Rate
Not Supported
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider Groq (Meta)
Model Family Llama 3.3
Blended 3:1 Rate (Production Benchmark) $0.640 / 1M tokens
Batch API Discount (24hr SLA) None
Multimodal Vision ❌ Text Only
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 86.0%
Latency & Throughput Tier Instant (~280 t/s)
Optimal Architecture & Use Cases Instantaneous interactive chatbots, real-time voice translation, low-latency tool agents

Frequently Asked Questions about Llama 3.3 70B (Groq LPU)

How much does Llama 3.3 70B (Groq LPU) cost per 1M tokens?

Llama 3.3 70B (Groq LPU) pricing is set at $0.590 per 1 million input tokens and $0.790 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a None.

How much money does Llama 3.3 70B (Groq LPU) prompt caching save?

Prompt caching is currently not natively offered for Llama 3.3 70B (Groq LPU) on direct serverless endpoints.

What is the context window limit of Llama 3.3 70B (Groq LPU)?

Llama 3.3 70B (Groq LPU) has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Direct Matchups with Llama 3.3 70B (Groq LPU)

See how Llama 3.3 70B (Groq LPU) compares against other leading frontier and open-weights models in cost and latency.