LLM Reference

DeepSeek R1 Distill Llama 8B

Released
2025-01-20
Last refreshed
2026-05-19
Status
Researched 106d ago
Open sourceCommercial use: permittedLong context

DeepSeek R1 Distill Llama 8B is worth evaluating for long context when its provider route and context window match the workload.

Use it for

  • Teams evaluating long context
  • Workloads that can use a 128k context window
  • Buyers comparing 2 tracked provider routes

Do not use it for

  • Vision or document-understanding workloads
  • Strict JSON or tool-calling flows
Specifications
Released
2025-01-20
Context
128k
Parameters
8B
Architecture
Decoder Only
Knowledge cutoff
2023-12
Specialization
general
Openness
Open source
License
MITOSI-approvedCommercial use: permitted
Weights
Unknown
Code
Unknown
Training
Multi-stage
Created by

Advancing artificial general intelligence (AGI).

Hangzhou, Zhejiang, China
Founded 2023
Website
Pricing
Output / 1M
$0.200
Input / 1M
$0.200

Cheapest of 2 routes · Fireworks AI

About

DeepSeek R1 Distill Llama 8B is DeepSeek's DeepSeek R1 model with an optional reasoning mode. It offers a 128K-token context window with weights openly available for self-hosting.

Top use-case fit

Long context

Included by capability and metadata signals in the decision map.

Capabilities

Reasoning

Benchmark peer barsfor Long context

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.