Kosmos-2 Models by Microsoft Research
Last refreshed 2026-05-19. Next refresh: weekly.
Details
About
The Kosmos-2 family of large language models (LLMs) is a significant advancement in the realm of multimodal AI, particularly noted for its ability to ground language understanding in real-world contexts. These models effectively integrate visual and textual information, excelling in tasks that involve perceiving object descriptions, such as identifying bounding boxes, and aligning textual data with visual content. Utilizing a Transformer-based architecture, Kosmos-2 is trained on a substantial dataset of grounded image-text pairs, enabling it to perform a range of tasks, including multimodal grounding, referring expression comprehension, and general language understanding and generation. Noteworthy is its innovative approach of representing referential expressions as Markdown links, which enhances the precision of visual-textual alignment. This positions the Kosmos-2 family as a vital bridge between language and multimodal perception, with its models like kosmos-2-patch14-224 available on Hugging Face, facilitating developments in areas such as image captioning and visual question answering.
Decision facts
Archived Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Keep only for existing workloads; choose a current variant for new builds.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| Kosmos 2 | Keep only for existing workloads; choose a current variant for new builds. | 2023-03 | 2k context1.7B parameters | Archived |
Release Timeline
1 release groupSpecifications(1 models)
| Model | Released | Context | Parameters |
|---|---|---|---|
| Kosmos 2 | 2023-03 | 2k | 1.66B |





