Last refreshed 2026-06-29. Next refresh: weekly.
Why use MOSS-Audio 4B Thinking on Hugging Face Inference Endpoints?
Hugging Face Inference Endpoints offers MOSS-Audio 4B Thinking with competitive pricing. Hugging Face is a leading AI community and platform dedicated to democratizing artificial intelligence.
Input / 1M
-
Output / 1M
-
Cache
Not sourced
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: OpenMOSS-Team/MOSS-Audio-4B-ThinkingModel ID
OpenMOSS-Team/MOSS-Audio-4B-ThinkingRequest example
Curated snippets for this provider are not sourced yet. Use Hugging Face Inference Endpoints documentation with model ID
OpenMOSS-Team/MOSS-Audio-4B-Thinking.Gotchas
- Use provider model ID "OpenMOSS-Team/MOSS-Audio-4B-Thinking", not the LLMReference slug "moss-audio-4b-thinking".
Capabilities
MultimodalReasoningAudio
About MOSS-Audio 4B Thinking
MOSS-Audio 4B Thinking is the reasoning-tuned 4.6B variant of MOSI AI and OpenMOSS Team's open-weight audio understanding model. It uses the MOSS-Audio encoder and Qwen3-4B backbone, adding chain-of-thought-oriented post-training for stronger complex audio reasoning while retaining speech, sound, music, timestamp, captioning, and QA coverage.
Get Started
Model Specs
Released2026-04-13
Parameters4.6B
ArchitectureAudio / Speech