LLM Reference
Hugging Face Inference Endpoints

Using MOSS-Audio 4B Thinking on Hugging Face Inference Endpoints

Implementation guide · MOSS-Audio · MOSI AI

Open Source

Hugging Face Inference Endpoints exposes MOSS-Audio 4B Thinking through model ID OpenMOSS-Team/MOSS-Audio-4B-Thinking. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-29. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at Hugging Face Inference Endpoints and generate an API key.
  2. 2
    Use the Hugging Face Inference Endpoints SDK or REST API to call OpenMOSS-Team/MOSS-Audio-4B-Thinking — see the documentation for request format.

Code Examples

Pricing on Hugging Face Inference Endpoints

Capabilities

MultimodalReasoningAudio

About MOSS-Audio 4B Thinking

MOSS-Audio 4B Thinking is the reasoning-tuned 4.6B variant of MOSI AI and OpenMOSS Team's open-weight audio understanding model. It uses the MOSS-Audio encoder and Qwen3-4B backbone, adding chain-of-thought-oriented post-training for stronger complex audio reasoning while retaining speech, sound, music, timestamp, captioning, and QA coverage.

Model Specs

Released2026-04-13
Parameters4.6B
ArchitectureAudio / Speech

Provider

Hugging Face Inference Endpoints
Hugging Face Inference Endpoints

Hugging Face

New York City, New York, United States