Nemotron-Labs TwoTower 30B-A3B Base
Nemotron-Labs TwoTower 30B-A3B Base is a released coding, long context, and classification model with open-weight and 128k context; evaluate it while provider pricing coverage matures.
- Family
- Nemotron-Labs TwoTower
- Released
- 2026-06-25
- Context
- 128k
- Max output
- 128,000
- Parameters
- ~60B total checkpoint; Hugging Face reports 63B params
- Architecture
- MoE + SSM Hybrid
- Knowledge cutoff
- 2025-06
- Specialization
- general
- Openness
- Open weights
- License
- NVIDIA Open ModelCommercial use: permitted
- Weights
- Available
- Code
- Unknown
- Training
- Pretrained
No tracked provider token pricing is available yet.
About
Base text-generation checkpoint for Nemotron-Labs TwoTower. It uses a two-tower block-diffusion architecture over a Mamba-2/Transformer hybrid MoE backbone: a frozen causal AR/context tower processes clean prompt and committed tokens, while a trainable diffusion/denoiser tower fills token blocks by mask diffusion with cross-attention to the context tower. The checkpoint ships both towers and is not an instruction-tuned model.
Provider price ladder
No tracked provider token pricing is available for this model yet.
Capabilities
No model capability flags are currently sourced.
Benchmark peer barsfor Coding
Benchmark scores(10)
| Benchmark | Score | Version | Evaluation | Source |
|---|---|---|---|---|
| Massive Multitask Language Understanding | 78.2 | 5-shot, accuracy; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| MMLU PRO | 60.9 | 5-shot, chain-of-thought exact match; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| AI2 Reasoning Challenge | 92.7 | ARC-Challenge, 25-shot, acc_norm; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| WinoGrande | 76.1 | 5-shot, accuracy; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| ReAding Comprehension Dataset From Examinations | 88.9 | 0-shot, accuracy; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| HumanEval | 75.6 | 0-shot; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| Mostly Basic Programming Problems | 74.3 | MBPP-Sanitized, 3-shot; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| Grade School Math 8K | 90.1 | 8-shot, accuracy; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| MATH-500 | 80.6 | 4-shot; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
| Multilingual Grade School Math | 80.4 | 8-shot, average accuracy; default TwoTower diffusion decoding at confidence_threshold=0.8, block_size=16, BF16 on 2xH100; evaluator/harness not published on the model cardObserved 2026-06-25 | — | Source |
Migration checks
No linked migration route is available for this model yet.
API versions
v1.1v1.0No tracked provider token pricing is available yet.