Last refreshed 2026-06-01. Next refresh: weekly.
Why use Cosmos 3 Nano on NVIDIA NIM?
NVIDIA NIM offers Cosmos 3 Nano with competitive pricing. NVIDIA NIM is NVIDIA's deployment platform for GPU-accelerated inference microservices.
Setup recipe
Docs fallbackUse the provider REST API or SDKCreate a provider API keymodel: cosmos3-reasoner-nanocosmos3-reasoner-nanoRequest example
cosmos3-reasoner-nano.Gotchas
- Use provider model ID "cosmos3-reasoner-nano", not the LLMReference slug "cosmos-3-nano".
Pricing
| Type | Price (per 1M) |
|---|---|
| Image input | $1.00 |
| Video input | $1.00 |
| Audio input | $1.00 |
Capabilities
About Cosmos 3 Nano
Cosmos 3 Nano is NVIDIA's 16B-parameter omnimodel optimized for efficient inference on workstation-grade hardware (NVIDIA RTX PRO 6000). Architecture: dual-tower Mixture-of-Transformers with an 8B autoregressive Reasoner and an 8B diffusion-based Generator. The Reasoner supports up to 256K tokens of context for vision-language reasoning; the Generator produces video up to 720p at variable frame rates (default 189 frames). Natively handles text, image, video, audio (48kHz stereo), and robot action trajectories across 10+ robot embodiments including Franka Panda, UR, Google robot, and UMI. BF16 precision only. Available as open weights on Hugging Face and via the Cosmos 3 Reasoner NIM (NIM_MODEL_SIZE=nano).