Skip to content
AI Primer

NVIDIA Nemotron is a family of open, multimodal AI models for long-running and agentic AI workflows, with open weights, training data, and recipes. The family includes reasoning models such as Nemotron 3 Nano, Super, and Ultra, plus models for visual understanding, speech, retrieval, and safety.

Model Intelligence

Benchmarkable
No
Model level
family

Recent stories

5 linked stories
workflowSECONDARY2026-07-16
AI21 reports 80.8% on SWE-Bench Pro with model-team agent

AI21 says its pipeline routes exploration, extraction, and patching across model tiers and reaches 80.8% on SWE-Bench Pro at $5.99 per task. Other workflows use critic agents, readable harnesses, Fugu/Nemotron routing, and scheduled Gemini agents.

releaseSECONDARY2026-07-12
German consortium releases SOOFI base model trained on 27T tokens

A German consortium released the small SOOFI sovereign base model trained on 27T tokens. Analysts said it reuses Nemotron 3 Nano architecture and many hyperparameters with a changed data mix, and benchmarks drew criticism for overstating capability versus Qwen and Nemotron.

releaseSECONDARY2026-06-04
NVIDIA releases Nemotron 3 Ultra: 550B MoE, 1M context

NVIDIA shipped Nemotron 3 Ultra, a 550B/55B-active hybrid Mamba-Transformer MoE with open weights, data, and recipe, plus broad runtime and host support. It matters because the model pairs frontier open benchmarks with immediate agent-serving options, though local use still needs heavy quantization or large-memory hardware.

releaseSECONDARY2026-04-28
Nemotron 3 Nano Omni launches 30B-A3B multimodal model with 256K context

NVIDIA opened Nemotron 3 Nano Omni, a 30B-A3B model for text, image, audio, and video, with day-one serving support. That lets teams run one open model for perception-heavy agents instead of stitching separate components.

releaseSECONDARY2026-03-13
NVIDIA releases Nemotron 3 Super: 120B open model targets 1M-token agent workloads

NVIDIA released Nemotron 3 Super, a 120B open model with 1M-token context and a hybrid architecture tuned for agent workloads, then landed it in Perplexity and Baseten. Try it if you need an open-weight long-context option that is already available in hosted stacks.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.