A new arXiv paper introduces Nemotron 3 Ultra, a 550 billion total parameter mixture-of-experts model with 55 billion active parameters. The hybrid Mamba-Transformer architecture is designed for efficient long-context and agentic reasoning workloads.

The model includes a 1 million-token context extension, supervised fine-tuning, reinforcement learning, distillation, and reasoning-budget control. The authors claim much higher inference throughput while maintaining competitive accuracy.

The release is another sign that frontier model design is diversifying beyond dense Transformers, especially where long context and agentic reasoning make efficiency critical.