NVIDIA Nemotron 3.5 Lightning is now available through Amazon SageMaker JumpStart, AWS says. The model is a 30-billion-parameter mixture-of-experts system with about 3 billion active parameters per request.

A mixture-of-experts model activates only parts of the network for a given input, which can reduce serving cost and latency compared with using every parameter every time. AWS describes Nemotron 3.5 Lightning as built for high-volume agentic workloads.

The post says the model can deliver up to four times higher throughput and up to 30 percent faster task completion for always-on agents. Those figures are vendor-reported and depend on deployment details, workload, and comparison baseline.

For developers already using SageMaker JumpStart, the update lowers the friction of trying the model in managed AWS infrastructure. The more important evaluation remains workload-specific: whether the model’s speed, quality, and cost fit a real agent application.