RecurrentGemma Models by Google DeepMind
Last refreshed 2026-05-19. Next refresh: weekly.
Details
About
RecurrentGemma is a family of open-weight language models developed by Google DeepMind, known for their cutting-edge Griffin architecture. This hybrid design blends linear recurrences with local attention mechanisms, allowing the models to excel in a range of language tasks with reduced memory overhead and efficient inference, especially on lengthy sequences. Unlike traditional transformer models that require memory scaling linearly with sequence length, RecurrentGemma maintains a fixed-sized state, resulting in faster processing speeds. Both pre-trained and instruction-tuned variants are available, the latter being tailored for tasks like dialogue and instruction following. Accessible through platforms like Hugging Face and Kaggle, RecurrentGemma-2B achieves performance akin to Gemma-2B despite being trained on fewer tokens, demonstrating its efficiency and versatility 23910.
Decision facts
- Best fit
- chatbot and role-playing use cases
- Capability starting point
- RecurrentGemma 9B with 4k context
- Lowest tracked input
- Not tracked
- Closest related family
- T5Gemma
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Use when the workload needs 4k context and 9B parameters.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| RecurrentGemma 9B | Use when the workload needs 4k context and 9B parameters. | 2024-06 | 4k context9B parameters | Current |
Release Timeline
2 release groupsSpecifications(2 models)
| Model | Released | Context | Parameters |
|---|---|---|---|
| RecurrentGemma 9B | 2024-06 | 4k | 9B |

