LLM Reference
Hugging Face Inference Endpoints

Using MOVA 720p on Hugging Face Inference Endpoints

Implementation guide · MOVA · MOSI AI

Open Source

Hugging Face Inference Endpoints exposes MOVA 720p through model ID OpenMOSS-Team/MOVA-720p. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-29. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Hugging Face Inference Endpoints and generate an API key.
  2. 2
    Use the Hugging Face Inference Endpoints SDK or REST API to call OpenMOSS-Team/MOVA-720p — see the documentation for request format.

Code Examples

Pricing on Hugging Face Inference Endpoints

Capabilities

VisionMultimodalAudio

About MOVA 720p

MOVA 720p is the higher-resolution open-weight MOVA checkpoint for synchronized video-audio generation. MOSI AI and the OpenMOSS Team describe MOVA as a 32B-parameter mixture-of-experts model with 18B active parameters during inference, designed for native image-to-video-audio and text-to-video-audio generation with synchronized audio, lip sync, and sound effects.

Model Specs

Released2026-01-29
Parameters32B total / 18B active
ArchitectureMixture of Experts

More Models on Hugging Face Inference Endpoints

Provider

Hugging Face Inference Endpoints
Hugging Face Inference Endpoints

Hugging Face

New York City, New York, United States