Home  /  Models  /  Llama 3.1 8B Instant (Groq)
⚡ Quick Pricing Summary

Llama 3.1 8B Instant (Groq) by Groq (Meta) costs $0.050 per 1 Million input tokens and $0.080 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $0.058/1M tokens.

Groq (Meta) Ultra-Fast / Low Cost

Llama 3.1 8B Instant (Groq)

Micro-footprint model hosted on LPU chips achieving 500+ tokens per second at near-zero cost.

Calculate Spend in App →
Input Token Price
$0.050
Per 1,000,000 tokens
Output Token Price
$0.080
Per 1,000,000 tokens
Prompt Caching Rate
Not Supported
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider Groq (Meta)
Model Family Llama 3.1
Blended 3:1 Rate (Production Benchmark) $0.058 / 1M tokens
Batch API Discount (24hr SLA) None
Multimodal Vision ❌ Text Only
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 73.0%
Latency & Throughput Tier Blazing (~550 t/s)
Optimal Architecture & Use Cases Instant search indexing, real-time moderation, intent classification

Frequently Asked Questions about Llama 3.1 8B Instant (Groq)

How much does Llama 3.1 8B Instant (Groq) cost per 1M tokens?

Llama 3.1 8B Instant (Groq) pricing is set at $0.050 per 1 million input tokens and $0.080 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a None.

How much money does Llama 3.1 8B Instant (Groq) prompt caching save?

Prompt caching is currently not natively offered for Llama 3.1 8B Instant (Groq) on direct serverless endpoints.

What is the context window limit of Llama 3.1 8B Instant (Groq)?

Llama 3.1 8B Instant (Groq) has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Direct Matchups with Llama 3.1 8B Instant (Groq)

See how Llama 3.1 8B Instant (Groq) compares against other leading frontier and open-weights models in cost and latency.