Llama 3 8B Instruct
Llama 3 8B Instruct is worth evaluating for coding, classification, and json / tool use when its provider route and context window match the workload.
Use it for
- Teams evaluating coding, classification, and json / tool use
- Workloads that can use a 8k context window
- Buyers comparing 4 tracked provider routes
Do not use it for
- Vision or document-understanding workloads
- Family
- Llama 3
- Released
- 2024-04-18
- Context
- 8k
- Parameters
- 8B
- Architecture
- Decoder Only
- Knowledge cutoff
- 2023-03
- Specialization
- general
- Openness
- Open weights
- License
- Llama 3 CommunityCommercial use: conditional
- Weights
- Available
- Code
- Unknown
- Training
- Fine-tuned
Large-scale open-source AI for social technologies.
Cheapest of 17 routes · OpenRouter
About
The Llama 3 8B Instruct model, released on April 18, 2024, is Meta's latest instruction-following language model with 8 billion parameters. It utilizes an auto-regressive transformer architecture with Grouped-Query Attention for improved scalability. Trained on over 15 trillion tokens and fine-tuned with 10 million human-annotated examples, it excels in dialogue and conversational tasks. The model outperforms its predecessors on industry benchmarks, scoring 68.4 on MMLU (5-shot). Designed for commercial and research applications, it prioritizes safety and helpfulness, making it suitable for chatbots, virtual assistants, and other interactive AI applications. For more details, visit the Hugging Face page [1].
Llama 3 8B Instruct is an open-weight model in the Llama 3 family. The structured metadata tracks a 8k-token context window and structured outputs. This page tracks provider routes through AWS Bedrock, DeepInfra, OctoAI API (Deprecated), and 14 more, with the cheapest tracked route listed at $0.02 input and $0.05 output per 1M tokens. Headline tracked benchmarks include Google-Proof Q&A 44.8, HellaSwag 91.1, and HumanEval 68.2.
Top use-case fit: coding, agents, and build tasks
Coding
Q/$ A1 relevant benchmark in the decision map.
Classification
Q/$ A3 relevant benchmarks in the decision map.
JSON / Tool use
Included by capability and metadata signals in the decision map.
Provider price ladder
Compare all 17Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Route |
|---|---|---|---|
| OpenRouter | $0.030 | $0.040 | Serverless |
| Novita AI | $0.040 | $0.040 | Serverless |
| DeepInfra | $0.020 | $0.050 | Serverless |
| Lepton AI API | $0.070 | $0.070 | Serverless |
Available via routers & gateways(16)
LiteLLM
GatewayOpen-source Python SDK and proxy server that unifies 100+ LLM APIs behind a single OpenAI-compatible interface, with load balancing, cost tracking, and configurable failover.
OpenRouter
HybridUnified hybrid gateway to 400+ models from 60+ providers via a single OpenAI-compatible API, with optional auto-routing that selects the best model per prompt.
Portkey
GatewayProduction AI gateway routing to 1,600+ LLMs with failover, load balancing, semantic caching, and guardrails; Apache 2.0 core is fully self-hostable with the complete feature set.
AIRouter
RouterCommercial LLM router that analyzes incoming requests and routes to the optimal model for cost/quality/latency via a drop-in OpenAI-compatible API, with a privacy-preserving embedding mode that avoids sending prompt content.
Amazon Bedrock Intelligent Prompt Routing
RouterAWS Bedrock's native intelligent prompt router that routes prompts between Anthropic Claude model tiers (Haiku/Sonnet) based on predicted task complexity, with no extra per-routing charge.
Azure AI Foundry Model Router
RouterMicrosoft Azure AI Foundry's native model router that uses a trained ML model to route each prompt in real time to the optimal Azure-hosted model, with Balanced/Cost/Quality mode selection and automatic failover.
Capabilities
Benchmark peer barsfor Coding
Benchmark scores(7)
| Benchmark | Score | Version | Source |
|---|---|---|---|
| Google-Proof Q&A | 44.8 | diamond | https://arxiv.org/abs/2407.21783 |
| HellaSwag | 91.1 | 10-shot | research |
| HumanEval | 68.2 | pass@1 | research |
| Massive Multitask Language Understanding | 76.9 | 5-shot | https://arxiv.org/abs/2407.21783 |
| Instruction-Following Evaluation | 59.5 | v2 | https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard |
| MMLU PRO | 40.5 | — | https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro |
| Grade School Math 8K | 80.6 | — | https://arxiv.org/abs/2407.21783 |
Migration checks
No linked migration route is available for this model yet.
Rankings & picks(3)
Compare Llama 3 8B Instruct with other models
- Llama 3 8B Instruct vs DeepSeek V4 Flash157
- Llama 3 8B Instruct vs Qwen3.6-35B-A3B150
- Llama 3 8B Instruct vs Phi-3 Mini 4k146
- Llama 3 8B Instruct vs Claude Opus 4.6144
- Llama 3 8B Instruct vs Llama 3.1 70B Instruct140
- Llama 3 8B Instruct vs DeepSeek V4 Pro134
- Llama 3 8B Instruct vs Qwen3-Max122
- Llama 3 8B Instruct vs Llama 3.2 1B Instruct119
Comparison and alternatives
Browse all comparisons →Show all 51 popular comparisonssorted by 7-day search impressions
Frequently asked questions
What is the context window of Llama 3 8B Instruct?
Llama 3 8B Instruct has a context window of 8k tokens.
How much does Llama 3 8B Instruct cost?
Llama 3 8B Instruct pricing ranges from $0.02/1M to $0.6/1M input tokens depending on the provider.
When was Llama 3 8B Instruct released?
Llama 3 8B Instruct was released on 2024-04-18.
Which providers offer Llama 3 8B Instruct?
Llama 3 8B Instruct is available from 17 providers: AWS Bedrock, DeepInfra, OctoAI API (Deprecated), Fireworks AI, Alibaba Cloud PAI-EAS, Baseten API, Lepton AI API, GCP Vertex AI, Cloudflare Workers AI, NVIDIA NIM, Together AI, Databricks Foundation Model Serving, IBM watsonx, Microsoft Foundry, OpenRouter, Replicate API, Novita AI.
What benchmarks has Llama 3 8B Instruct been tested on?
Llama 3 8B Instruct has been evaluated on 7 benchmarks, including Google-Proof Q&A, HellaSwag, HumanEval, Massive Multitask Language Understanding, Instruction-Following Evaluation.
Large-scale open-source AI for social technologies.
Cheapest of 17 routes · OpenRouter