Fuyu Models by Adept AI
Last refreshed 2026-05-19. Next refresh: weekly.
Details
About
Developed by Adept AI, the Fuyu large language model (LLM) family is distinct for its innovative approach in handling multimodal tasks, specifically catering to digital agents. It boasts a simplified architecture, employing a decoder-only transformer without needing a specific image encoder. Instead, image patches are directly fed as linear projections into the transformer's initial layer, which allows for flexible processing of various image resolutions. This configuration not only streamlines training but also accelerates inference, providing rapid response times under 100 milliseconds for handling large images. Although the Fuyu models are primarily optimized for digital agent applications, they demonstrate robust performance on standard image understanding benchmarks. One of the models in this family, the Fuyu-8B, is publicly available under a CC-BY-NC license, offering opportunities for research and development with the caveat that it may require fine-tuning to meet specific application needs.
Decision facts
- Best fit
- codingagent workflows
- Capability starting point
- Fuyu-Heavy
- Lowest tracked input
- Not tracked
- Closest related family
- Persimmon
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| Fuyu-Heavy | Use when provider availability and model metadata match the workload. | 2024-01 | — | Current |
Release Timeline
1 release groupSpecifications(2 models)
| Model | Released | Parameters |
|---|---|---|
| Fuyu-Heavy | 2024-01 | — |

