Last refreshed 2026-06-29. Next refresh: weekly.
Why use MOSS-Audio 4B Instruct on Hugging Face Inference Endpoints?
Hugging Face Inference Endpoints offers MOSS-Audio 4B Instruct with competitive pricing. Hugging Face is a leading AI community and platform dedicated to democratizing artificial intelligence.
Input / 1M
-
Output / 1M
-
Cache
Not sourced
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: OpenMOSS-Team/MOSS-Audio-4B-InstructModel ID
OpenMOSS-Team/MOSS-Audio-4B-InstructRequest example
Curated snippets for this provider are not sourced yet. Use Hugging Face Inference Endpoints documentation with model ID
OpenMOSS-Team/MOSS-Audio-4B-Instruct.Gotchas
- Use provider model ID "OpenMOSS-Team/MOSS-Audio-4B-Instruct", not the LLMReference slug "moss-audio-4b-instruct".
Capabilities
MultimodalAudio
About MOSS-Audio 4B Instruct
MOSS-Audio 4B Instruct is the instruction-following 4.6B variant of MOSI AI and OpenMOSS Team's open-weight audio understanding model. It combines a MOSS-Audio encoder with a Qwen3-4B language backbone for speech, environmental sound, music, captioning, time-aware question answering, timestamped ASR, and audio-grounded reasoning.
Get Started
Model Specs
Released2026-04-13
Parameters4.6B
ArchitectureAudio / Speech