LLM Reference

Llama 3.1 Models by AI at Meta

AI at MetaLlama 3 CommunityOpen weightsHighlightOpen Source
8 models2024Up to 128k ctxFrom $0.02/1M input

Details

ResearcherAI at Meta
Commercial useCommercial use: conditional
Models8
Released2024
Max context128k

Capabilities

Function Calling2 of 8 models
Tool Use2 of 8 models
Structured Outputs4 of 8 models

About

The Llama 3.1 family, developed by Meta, features large language models (LLMs) with sizes of 8B, 70B, and 405B parameters 12. The standout 405B model is the largest openly available foundation model, showcasing capabilities that compete with top closed-source models 1. These models boast a 128K token context window, support for eight languages, and enhanced reasoning capabilities, making them versatile for multiple applications. The 405B model is ideal for synthetic data generation and model distillation, highlighting Meta's focus on open-source AI, enabling customization and deployment without needing to share data with Meta 1.

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

8 in view

Use when the workload needs 128k context, 405B parameters, and structured outputs.

2024-07128k context405B parametersstructured outputs

Use when the workload needs 128k context, 70B parameters, and structured outputs.

2024-07128k context70B parametersstructured outputs

Use when the workload needs 128k context, 8B parameters, and structured outputs.

2024-07128k context8B parametersstructured outputs

Use when the workload needs 128k context and 405B parameters.

2024-07128k context405B parameters

Use when the workload needs 128k context and 70B parameters.

2024-07128k context70B parameters

Use when the workload needs 128k context and 8B parameters.

2024-07128k context8B parameters

Use when the workload needs 128k context, 405B parameters, and tool use.

2024-07128k context405B parameterstool use

Use when the workload needs 128k context, 70B parameters, and tool use.

2024-07128k context70B parameterstool use

Release Timeline

1 release group
2024-07
8 current
Llama 3.1 405B
128k context405B parameters
Current
Llama 3.1 405B Instruct
128k context405B parametersstructured outputs
Current
Llama 3.1 70B
128k context70B parameters
Current
Llama 3.1 70B Instruct
128k context70B parametersstructured outputs
Current
Llama 3.1 8B
128k context8B parameters
Current
Llama 3.1 8B Instruct
128k context8B parametersstructured outputs
Current
Llama 3.1-405B
128k context405B parameterstool use
Current
Llama 3.1-70B
128k context70B parameterstool use
Current

Specifications(8 models)

Llama 3.1 model specifications comparison
ModelReleasedContextParametersFn CallingTool UseStructured Outputs
Llama 3.1 405B Instruct2024-07128k405BNoNoYes
Llama 3.1 70B Instruct2024-07128k70BNoNoYes
Llama 3.1 8B Instruct2024-07128k8BNoNoYes
Llama 3.1 405B2024-07128k405BNoNoNo
Llama 3.1 70B2024-07128k70BNoNoNo
Llama 3.1 8B2024-07128k8BNoNoNo
Llama 3.1-405B2024-07128k405BYesYesYes
Llama 3.1-70B2024-07128k70BYesYesNo

Pricing

Llama 3.1 model pricing by provider
ModelProviderInput / 1MOutput / 1MType
Llama 3.1 8B InstructOpenRouter$0.02$0.05Serverless
Llama 3.1 8B InstructNovita AI$0.02$0.05Serverless
Llama 3.1 8B InstructDeepInfra$0.02$0.05Serverless
Llama 3.1 8B InstructGroqCloud$0.05$0.08Serverless
Llama 3.1 8B InstructHyperbolic AI Inference$0.1$0.1Serverless
Llama 3.1 8B InstructIBM watsonx$0.15$0.5Serverless
Llama 3.1 8B InstructTogether AI$0.18$0.18Serverless
Llama 3.1 8B InstructFireworks AI$0.2$0.2Serverless
Llama 3.1 8B InstructAWS Bedrock$0.22$0.22Serverless
Llama 3.1 8B InstructVercel AI Gateway$0.22$0.22Serverless
Llama 3.1 8B InstructReplicate API$0.25$0.25Serverless
Llama 3.1 8B InstructMicrosoft Foundry$0.3$0.61Provisioned
Llama 3.1 70B InstructHyperbolic AI Inference$0.4$0.4Serverless
Llama 3.1 70B InstructDeepInfra$0.4$0.4Serverless
Llama 3.1 70B InstructOpenRouter$0.4$0.4Serverless
Llama 3.1 70B InstructAWS Bedrock$0.72$0.72Serverless
Llama 3.1 70B InstructVercel AI Gateway$0.72$0.72Serverless
Llama 3.1 70B InstructIBM watsonx$0.8$2.4Serverless
Llama 3.1 70B InstructTogether AI$0.88$0.88Serverless
Llama 3.1 70B InstructFireworks AI$0.9$0.9Serverless
Llama 3.1-70BReplicate API$1.2$1.2Serverless
Llama 3.1 405B InstructAWS Bedrock$2.4$2.4Serverless
Llama 3.1 70B InstructMicrosoft Foundry$2.68$3.54Provisioned
Llama 3.1 405B InstructFireworks AI$3$3Serverless
Llama 3.1 405B InstructIBM watsonx$3$9Serverless
Llama 3.1-405BReplicate API$3.75$3.75Serverless
Llama 3.1 405B InstructHyperbolic AI Inference$4$4Serverless
Llama 3.1 405B InstructTogether AI$5$15Serverless
Llama 3.1-405BGCP Vertex AI$5$16Serverless
Llama 3.1 405B InstructGCP Vertex AI$5$16Serverless
Llama 3.1 405B InstructMicrosoft Foundry$5.33$16Provisioned

Popular comparisons in this family

Frequently Asked Questions

What is Llama 3.1 used for?
Llama 3.1 is used for agent workflows and tool use, structured outputs, and coding. The family description and listed model capabilities point to those workloads as the best fit.
How does Llama 3.1 compare to MOSS-Audio?
Llama 3.1 by AI at Meta is strongest where you need agent workflows and tool use, while MOSS-Audio by MOSI AI is the closest related family to check for multimodal. Llama 3.1 has 8 listed variants and reaches up to 128k context, so compare the specs and pricing tables before choosing a production model.
Which Llama 3.1 model should I use?
For the lowest listed input price, start with Llama 3.1 8B Instruct through DeepInfra at $0.02/1M input tokens. For the most capable/latest local choice, evaluate Llama 3.1-405B with 128k context and tool use, function calling, and structured outputs.