LLM Reference

InternLM-XComposer2 Models by Intern-AI

Intern-AIApache 2.0Open source
3 models2024

Last refreshed 2026-04-19. Next refresh: weekly.

Details

ResearcherIntern-AI
LicenseApache 2.0OSI-approved
Commercial useCommercial use: permitted
Models3
Released2024

About

The InternLM-XComposer2 family brings together a suite of vision-language large models (VLLMs) based on the InternLM2 foundation model. These models are adept at tackling advanced text-image comprehension and composition tasks, with variations such as the InternLM-XComposer2-VL excelling in multimodal benchmarks and InternLM-XComposer2 tailored for sophisticated text-image composition. A standout in this series, the InternLM-XComposer2.5, offers capabilities comparable to GPT-4V with just a 7B LLM backend, while the InternLM-XComposer2-4KHD can comprehend images up to 4K resolution. Open-source and accessible on Hugging Face and GitHub, these models cater to various needs, including multimodal content creation and enhancing visual language understanding, with some versions optimized for devices with limited resources through 4-bit quantization.

Decision facts

Best fit
coding
Capability starting point
InternLM XComposer2 4KHD 7B
Lowest tracked input
Not tracked
Closest related family
InternVL

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

3 in view

Use when the workload needs 7B parameters.

2024-047B parameters

Use when the workload needs 7B parameters.

2024-047B parameters

Use when the workload needs 7B parameters.

2024-047B parameters

Release Timeline

1 release group
2024-04
3 current
Current
Current
Current

Specifications(3 models)

InternLM-XComposer2 model specifications comparison
ModelReleasedParameters
InternLM XComposer2 4KHD 7B2024-047B
InternLM XComposer2 7B2024-047B
InternLM XComposer2 VL 7B2024-047B