LLM Reference
NVIDIA NIM

Using Cosmos 3 Nano on NVIDIA NIM

Implementation guide · Cosmos 3 · NVIDIA AI

ProvisionedOpen Weights

NVIDIA NIM exposes Cosmos 3 Nano through model ID cosmos3-reasoner-nano. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-01. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at NVIDIA NIM and generate an API key.
  2. 2
    Use the NVIDIA NIM SDK or REST API to call cosmos3-reasoner-nano — see the documentation for request format.

Code Examples

See NVIDIA NIM documentation for integration details.

Pricing on NVIDIA NIM

TypePrice (per 1M)
Image input$1.00
Video input$1.00
Audio input$1.00

Capabilities

VisionMultimodalReasoningAudio

About Cosmos 3 Nano

Cosmos 3 Nano is NVIDIA's 16B-parameter omnimodel optimized for efficient inference on workstation-grade hardware (NVIDIA RTX PRO 6000). Architecture: dual-tower Mixture-of-Transformers with an 8B autoregressive Reasoner and an 8B diffusion-based Generator. The Reasoner supports up to 256K tokens of context for vision-language reasoning; the Generator produces video up to 720p at variable frame rates (default 189 frames). Natively handles text, image, video, audio (48kHz stereo), and robot action trajectories across 10+ robot embodiments including Franka Panda, UR, Google robot, and UMI. BF16 precision only. Available as open weights on Hugging Face and via the Cosmos 3 Reasoner NIM (NIM_MODEL_SIZE=nano).

Model Specs

Released2026-05-31
Parameters16B
Context256k
ArchitectureMixture of Transformers

Provider

NVIDIA NIM
NVIDIA NIM

NVIDIA

Santa Clara, California, United States