LLM Reference

Nemotron 3 Ultra

Released
2026-06-04
Last refreshed
2026-06-15
Status
Researched 85d ago
Open weightsCommercial use: permittedLong context

Nemotron 3 Ultra is worth evaluating for long context when its provider route and context window match the workload.

Use it for

  • Teams evaluating long context
  • Workloads that can use a 1m context window
  • Buyers comparing 1 tracked provider route

Do not use it for

  • Vision or document-understanding workloads
  • Strict JSON or tool-calling flows
Specifications
Released
2026-06-04
Context
1m
Parameters
550B
Architecture
Mixture of Experts
Specialization
general
Openness
Open weights
License
NVIDIA Open ModelCommercial use: permitted
Weights
Available
Code
Unknown
Created by

Accelerated AI for enterprise solutions

Santa Clara, California, United States
Founded 2015
Website
Pricing
Output / 1M
$2.20
Input / 1M
$0.500

Cheapest of 1 route · OpenRouter

About

NVIDIA's open frontier-reasoning model (550B total / 55B active MoE, hybrid Transformer-Mamba). Highest Artificial Analysis Intelligence Index for any US open model (score: 48). 300+ tokens/second. 1M-token context. Announced at Computex 2026. Pricing: ~$0.60/$2.60 per 1M tokens (provider median); free tier on some providers.

Top use-case fit

Long context

Included by capability and metadata signals in the decision map.

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
OpenRouter$0.500$2.20
Serverless

Capabilities

Reasoning

Benchmark peer barsfor Long context

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.