LLM Reference

Florence 2 Models by Microsoft Research

Microsoft ResearchMITOpen sourceOpen Source
2 models2024

Last refreshed 2026-04-15. Next refresh: weekly.

Details

LicenseMITOSI-approved
Commercial useCommercial use: permitted
Models2
Released2024

About

The Florence-2 family, created by Microsoft, features advanced large language models designed specifically for vision and vision-language tasks. These models are known for their ability to effectively address a variety of assignments, such as captioning, object detection, and segmentation, by employing a prompt-based methodology 2. Their unified representation is a significant advantage, allowing seamless task execution within a single model framework 3. Leveraging the extensive FLD-5B dataset, which offers 5.4 billion annotations across 126 million images, these models excel in multitask learning 2. The Florence-2 suite includes the Florence-2-base and Florence-2-large models, featuring parameter counts of 0.23 billion and 0.77 billion, respectively. Additionally, fine-tuned iterations like Florence-2-base-ft and Florence-2-large-ft demonstrate enhanced performance across various downstream tasks, while their compact size ensures they are efficient and suitable for resource-limited environments 3.

Decision facts

Best fit
coding
Capability starting point
Florence 2 Base
Lowest tracked input
Not tracked
Closest related family
Harrier

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

2 in view

Use when the workload needs 230M parameters.

2024-06230M parameters

Use when the workload needs 770M parameters.

2024-06770M parameters

Release Timeline

1 release group
2024-06
2 current
Florence 2 Base
230M parameters
Current
Florence 2 Large
770M parameters
Current

Specifications(2 models)

Florence 2 model specifications comparison
ModelReleasedParameters
Florence 2 Base2024-06230M
Florence 2 Large2024-06770M