Llama 4 Maverick 17B Instruct FP8
Llama 4 Maverick 17B Instruct FP8 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.
Use it for
- Teams evaluating coding, rag, and agents
- Workloads that can use a 1m context window
- Buyers comparing 4 tracked provider routes
Do not use it for
- Workloads where another current model has stronger sourced task evidence
- Family
- Llama 4
- Released
- 2025-04-05
- Context
- 1m
- Parameters
- 400B (17B active)
- Architecture
- Mixture of Experts
- Knowledge cutoff
- 2024-08
- Specialization
- general
- Openness
- Open weights
- License
- Llama 4 CommunityCommercial use: conditional
- Weights
- Unknown
- Code
- Unknown
- Training
- Fine-tuned
Large-scale open-source AI for social technologies.
Cheapest of 11 routes · DeepInfra
About
Meta's Llama 4 Maverick 17B with 128 experts, FP8-optimized for cost-efficient inference. Supports native Model Router integration on Microsoft Foundry.
Llama 4 Maverick 17B Instruct FP8 is an open-weight model in the Llama 4 family. The structured metadata tracks a 1m-token context window, multimodal input, and structured outputs. This page tracks provider routes through Microsoft Foundry, Together AI, OpenRouter, and 8 more, with the cheapest tracked route listed at $0.15 input and $0.6 output per 1M tokens. Headline tracked benchmarks include HumanEval 77.4, Aider Polyglot 15.6, and BigCodeBench 49.7.
Top use-case fit: coding, agents, and build tasks
Coding
Q/$ B4 relevant benchmarks in the decision map.
RAG
Included by capability and metadata signals in the decision map.
Agents
Q/$ A1 relevant benchmark in the decision map.
Provider price ladder
Compare all 11Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Route |
|---|---|---|---|
| DeepInfra | $0.150 | $0.600 | Serverless |
| OpenRouter | $0.150 | $0.600 | Serverless |
| Novita AI | $0.270 | $0.850 | Serverless |
| Together AI | $0.270 | $0.850 | Serverless |
Available via routers & gateways(16)
LiteLLM
GatewayOpen-source Python SDK and proxy server that unifies 100+ LLM APIs behind a single OpenAI-compatible interface, with load balancing, cost tracking, and configurable failover.
OpenRouter
HybridUnified hybrid gateway to 400+ models from 60+ providers via a single OpenAI-compatible API, with optional auto-routing that selects the best model per prompt.
Portkey
GatewayProduction AI gateway routing to 1,600+ LLMs with failover, load balancing, semantic caching, and guardrails; Apache 2.0 core is fully self-hostable with the complete feature set.
AIRouter
RouterCommercial LLM router that analyzes incoming requests and routes to the optimal model for cost/quality/latency via a drop-in OpenAI-compatible API, with a privacy-preserving embedding mode that avoids sending prompt content.
Amazon Bedrock Intelligent Prompt Routing
RouterAWS Bedrock's native intelligent prompt router that routes prompts between Anthropic Claude model tiers (Haiku/Sonnet) based on predicted task complexity, with no extra per-routing charge.
Azure AI Foundry Model Router
RouterMicrosoft Azure AI Foundry's native model router that uses a trained ML model to route each prompt in real time to the optimal Azure-hosted model, with Balanced/Cost/Quality mode selection and automatic failover.
Capabilities
Benchmark peer barsfor Coding
Benchmark scores(10)
| Benchmark | Score | Version | Source |
|---|---|---|---|
| HumanEval | 77.4 | 2025-04 | https://ai.meta.com/blog/llama-4-multimodal-intelligence/ |
| Aider Polyglot | 15.6 | 2026-04 | https://aider.chat/docs/leaderboards |
| BigCodeBench | 49.7 | 2025-04 (Instruct Pass@1) | https://bigcode-bench.github.io/results.json |
| Chatbot Arena | 1365.0 | — | https://lmarena.ai |
| Google-Proof Q&A | 67.1 | diamond | https://artificialanalysis.ai/leaderboards/models |
| τ-bench | 68.5 | τ-bench | https://taubench.com/ |
| MMMU Pro | 59.6 | LLM-Stats aggregator | https://llm-stats.com/benchmarks/mmmu-pro |
| Massive Multi-discipline Multimodal Understanding | 73.4 | — | https://ai.meta.com/blog/llama-4-scout-maverick/ |
| MMLU PRO | 80.5 | — | https://ai.meta.com/blog/llama-4-scout-maverick/ |
| LiveCodeBench | 43.4 | — | https://ai.meta.com/blog/llama-4-scout-maverick/ |
Migration checks
No linked migration route is available for this model yet.
Rankings & picks(2)
Compare Llama 4 Maverick 17B Instruct FP8 with other models
- Llama 4 Maverick 17B Instruct FP8 vs Llama 3.3 70B297
- Llama 4 Maverick 17B Instruct FP8 vs Llama 4 Scout 17B-16E Instruct124
- Llama 4 Maverick 17B Instruct FP8 vs DeepSeek V4 Pro33
- Llama 4 Maverick 17B Instruct FP8 vs Qwen3.6-Max20
- Llama 4 Maverick 17B Instruct FP8 vs Grok 414
- Llama 4 Maverick 17B Instruct FP8 vs Claude Sonnet 4.614
- Llama 4 Maverick 17B Instruct FP8 vs Mistral Large 28
- Llama 4 Maverick 17B Instruct FP8 vs DeepSeek V34
Comparison and alternatives
Browse all comparisons →Frequently asked questions
What is the context window of Llama 4 Maverick 17B Instruct FP8?
Llama 4 Maverick 17B Instruct FP8 has a context window of 1m tokens.
How much does Llama 4 Maverick 17B Instruct FP8 cost?
Llama 4 Maverick 17B Instruct FP8 pricing ranges from $0.15/1M to $0.35/1M input tokens depending on the provider.
When was Llama 4 Maverick 17B Instruct FP8 released?
Llama 4 Maverick 17B Instruct FP8 was released on 2025-04-05.
Which providers offer Llama 4 Maverick 17B Instruct FP8?
Llama 4 Maverick 17B Instruct FP8 is available from 11 providers: Microsoft Foundry, Together AI, OpenRouter, Fireworks AI, DeepInfra, GCP Vertex AI, NVIDIA NIM, AWS Bedrock, Vercel AI Gateway, Novita AI, Inceptron.
What benchmarks has Llama 4 Maverick 17B Instruct FP8 been tested on?
Llama 4 Maverick 17B Instruct FP8 has been evaluated on 10 benchmarks, including HumanEval, Aider Polyglot, BigCodeBench, Chatbot Arena, Google-Proof Q&A.
Large-scale open-source AI for social technologies.
Cheapest of 11 routes · DeepInfra