Cerebras LLaVA Models by Cerebras
Last refreshed 2026-06-15. Next refresh: weekly.
Details
About
Cerebras Systems' large language model family, based on the LLaVA architecture, are cutting-edge multimodal models capable of processing both text and images. These models, freely available on Hugging Face, come in various sizes, notably the 7B and 13B parameter versions 48. They are trained using Cerebras’s unique Wafer-Scale Engine, ensuring efficient and robust training processes. Enhancing their functionality, a vision encoder, such as the CLIP-VisionModel-Large, is integrated, empowering these models with visual instruction following and chat capabilities in a multimodal context. These models leverage Vicuna checkpoints for pretraining and are extensively fine-tuned on varied datasets, optimizing their performance for research in large multimodal models and chatbot applications. Notably, vision encoder checkpoints are provided separately, broadening their applicability in diverse projects 4.
Decision facts
- Best fit
- codingchatbot and role-playing use cases
- Capability starting point
- Cerebras LLaVA 13B with 4k context
- Lowest tracked input
- Not tracked
- Closest related family
- Cerebras GPT
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Use when the workload needs 4k context and 13B parameters.
Use when the workload needs 4k context and 7B parameters.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| Cerebras LLaVA 13B | Use when the workload needs 4k context and 13B parameters. | 2024-08 | 4k context13B parameters | Current |
| Cerebras LLaVA 7B | Use when the workload needs 4k context and 7B parameters. | 2024-08 | 4k context7B parameters | Current |
Release Timeline
1 release groupSpecifications(2 models)
| Model | Released | Context | Parameters |
|---|---|---|---|
| Cerebras LLaVA 13B | 2024-08 | 4k | 13B |
| Cerebras LLaVA 7B | 2024-08 | 4k | 7B |

