NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
Frontier accuracy, 5X greater speed, 30% lower cost.
A NVIDIA Nemotron 3 Ultra language-model release with 550B total parameters and 55B active parameters, using the NVFP4 checkpoint format. It is a hybrid Mamba-Transformer mixture-of-experts model release intended for long-context and agentic workloads.
Pricing
Model Intelligence
Recent stories
NVIDIA shipped Nemotron 3 Ultra, a 550B/55B-active hybrid Mamba-Transformer MoE with open weights, data, and recipe, plus broad runtime and host support. It matters because the model pairs frontier open benchmarks with immediate agent-serving options, though local use still needs heavy quantization or large-memory hardware.
Arena shipped Agent Mode, a benchmark that lets models use web search, bash, file writing, image generation, and follow-up questions, then ranks them on five live-session signals. It matters because agent evals move from static task sets to real user workflows, with GPT-5.5 High currently leading the leaderboard.
NVIDIA teased Nemotron 3 Ultra as a 550B open-weight model due later this week, with early messaging centered on 5x faster and 30% cheaper inference plus a hybrid SSM-MoE design. The rollout matters because early benchmark posts already place it near the top of open-weight leaderboards, widening NVIDIA’s open-model push beyond Cosmos.