LLM Reference

Cerebras GPT 256M

Released
2023-03-13
Last refreshed
2026-06-15
Status
Researched 103d ago
Open sourceCommercial use: permitted

Cerebras GPT 256M is released 2023-03-13 in the Cerebras GPT family with open-source; evaluate it while provider pricing coverage matures.

Use it for

  • Teams evaluating general LLM work
  • Workloads that can use a 2k context window

Do not use it for

  • Cost-sensitive launches that need sourced token pricing
  • Vision or document-understanding workloads
  • Strict JSON or tool-calling flows
Specifications
Released
2023-03-13
Context
2k
Parameters
256M
Architecture
Decoder Only
Knowledge cutoff
2020
Specialization
general
Openness
Open source
License
Apache 2.0OSI-approvedCommercial use: permitted
Weights
Unknown
Code
Unknown
Training
Fine-tuned
Created by

World's largest AI chip innovation

Sunnyvale, California, United States
Founded 2016
Website
Pricing

No tracked provider token pricing is available yet.

About

The Cerebras GPT 256M is a transformer-based large language model developed by Cerebras Systems, featuring a GPT-3 style architecture with 256 million parameters. It is part of a wider model family trained for compute-optimal performance according to Chinchilla scaling laws. It supports a vocabulary of 50,257 tokens and can handle sequences up to 2048 tokens long. Built for research, it demonstrates capabilities in text generation and language understanding, with potential for fine-tuning for conversational dialogue, though limited by its lack of instruction tuning.

Top use-case fit

No primary decision-task fit is mapped for this model yet.

Provider price ladder

No tracked provider token pricing is available for this model yet.

Capabilities

No model capability flags are currently sourced.

Benchmark peer barsfor Coding

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.