MOVA Models by MOSI AI
2 models2026
Last refreshed 2026-06-29. Next refresh: weekly.
Details
ResearcherMOSI AI
LicenseApache 2.0OSI-approved
Commercial useCommercial use: permitted
Models2
Released2026
Capabilities
VisionAll models
MultimodalAll models
About
MOVA is an open-weight video-audio generation family from MOSI AI and the OpenMOSS Team. It targets synchronized image-to-video-audio and text-to-video-audio generation with native audio, lip sync, sound effects, and an asymmetric dual-tower mixture-of-experts architecture.
Decision facts
- Best fit
- multimodalvideo generationvision and multimodal work
- Capability starting point
- MOVA 360p with multimodal inputs
- Lowest tracked input
- Not tracked
- Closest related family
- MOSS-Audio
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
2 in view
MOVA 360pCurrent
Use when the workload needs video generation, multimodal inputs, and audio.
2026-01video generationmultimodal inputsaudio
MOVA 720pCurrent
Use when the workload needs video generation, multimodal inputs, and audio.
2026-01video generationmultimodal inputsaudio
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| MOVA 360p | Use when the workload needs video generation, multimodal inputs, and audio. | 2026-01 | video generationmultimodal inputsaudio | Current |
| MOVA 720p | Use when the workload needs video generation, multimodal inputs, and audio. | 2026-01 | video generationmultimodal inputsaudio | Current |
