Llama 3.1 405B (Together AI) by Together AI (Meta) costs $3.50 per 1 Million input tokens and $3.50 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $3.500/1M tokens.
Llama 3.1 405B (Together AI)
The world's largest open-weights frontier model with 405 Billion parameters.
Verified Specifications & Benchmark Data
| Developer / Provider | Together AI (Meta) |
| Model Family | Llama 3.1 |
| Blended 3:1 Rate (Production Benchmark) | $3.500 / 1M tokens |
| Batch API Discount (24hr SLA) | None |
| Multimodal Vision | ❌ Text Only |
| Function Calling / Structured Outputs | ✅ Native Tool Calling |
| MMLU Benchmark Score | 88.6% |
| Latency & Throughput Tier | Standard (~30 t/s) |
| Optimal Architecture & Use Cases | Synthetic data generation, model distillation, enterprise on-premise benchmarking |
Frequently Asked Questions about Llama 3.1 405B (Together AI)
How much does Llama 3.1 405B (Together AI) cost per 1M tokens?
Llama 3.1 405B (Together AI) pricing is set at $3.50 per 1 million input tokens and $3.50 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a None.
How much money does Llama 3.1 405B (Together AI) prompt caching save?
Prompt caching is currently not natively offered for Llama 3.1 405B (Together AI) on direct serverless endpoints.
What is the context window limit of Llama 3.1 405B (Together AI)?
Llama 3.1 405B (Together AI) has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.
Direct Matchups with Llama 3.1 405B (Together AI)
See how Llama 3.1 405B (Together AI) compares against other leading frontier and open-weights models in cost and latency.
Compare Together AI (Meta) vs OpenAI pricing, context limits, and cost per 1M tokens.
Compare Together AI (Meta) vs Cohere pricing, context limits, and cost per 1M tokens.
Compare Together AI (Meta) vs Anthropic pricing, context limits, and cost per 1M tokens.