Using MOSS-Audio 4B Instruct on Hugging Face Inference Endpoints
Implementation guide · MOSS-Audio · MOSI AI
Open Source
Hugging Face Inference Endpoints exposes MOSS-Audio 4B Instruct through model ID OpenMOSS-Team/MOSS-Audio-4B-Instruct. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-06-29. Next refresh: weekly.
Quick Start
- 1
- 2Use the Hugging Face Inference Endpoints SDK or REST API to call
OpenMOSS-Team/MOSS-Audio-4B-Instruct— see the documentation for request format.
Code Examples
Pricing on Hugging Face Inference Endpoints
Capabilities
MultimodalAudio
About MOSS-Audio 4B Instruct
MOSS-Audio 4B Instruct is the instruction-following 4.6B variant of MOSI AI and OpenMOSS Team's open-weight audio understanding model. It combines a MOSS-Audio encoder with a Qwen3-4B language backbone for speech, environmental sound, music, captioning, time-aware question answering, timestamped ASR, and audio-grounded reasoning.
Model Specs
Released2026-04-13
Parameters4.6B
ArchitectureAudio / Speech