SemiAnalysis reports up to 10× better inference performance per dollar on Rubin NVL72
SemiAnalysis reports up to 10× better inference performance per dollar on Rubin NVL72 than GB300 NVL72 using vLLM. These are preview results; the full public InferenceX submission is still pending.

TL;DR
- Rubin NVL72 reaches up to 10× inference performance per dollar and 3.2× modeled profit per gigawatt versus GB300 NVL72 using vLLM, according to SemiAnalysis’s comparison.
- The results remain an InferenceX preview, with the full public submission still forthcoming, as SemiAnalysis clarified in a follow-up.
- MiniMax M3 delivered up to 7.84× GB200’s per-chip throughput at matched interactivity in vLLM’s early benchmarks.
- Rubin support is available through
vllm/vllm-openai:cu134-nightly, with DeepSeek, Kimi, GLM, and MiniMax coverage described in vLLM’s announcement.
A nice NUMA trick: vLLM’s locality-aware MoE runs faster even when memory-domain partitioning leaves 12 of 212 SMs unused. SGLang fused four MoE tail operations into one collective kernel, removing 276 launches per decode step.
Performance per dollar and profit
MiniMax M3 428B reaches $39 billion in modeled annual profit in SemiAnalysis’s estimate, with the MiniMax calculator expressing economics per gigawatt of utility power.
Annual profit means token revenue minus modeled serving costs and applicable model-license fees under InferenceX’s definition. The calculation depends on:
- Token prices and cache mix.
- Target interactivity and utilization.
- Owning versus renting compute, plus any license fees.
Costs outside the model can change realized profit, and the definition explicitly distinguishes the estimate from audited corporate net income.
AgentX and MLPerf baselines
vLLM reports three throughput comparisons in its benchmark write-up:
- AgentX, MiniMax M3 428B: 5.18× GB200 NVL72’s per-chip throughput at a P90 interactivity target of 150 tokens/s/user.
- AgentX, the same model: 7.84× GB200’s per-chip throughput at approximately 440 tokens/s/user.
- MLPerf Inference v6.1, Qwen3-VL-235B-A22B: up to 3.7× GB300 NVL72’s throughput across offline, server, and interactive scenarios, using vLLM as the inference backend and Dynamo as the frontend router.
MiniMax M3 per-chip throughput versus P90 interactivity, marked as an InferenceX Official Preview.
InferenceX preview status
The full public vLLM Rubin NVL72 submission and a SGLang Rubin NVL72 submission are still forthcoming. vLLM says it will broaden testing as it gains access to more NVL72 nodes in the announcement.
Rubin hardware and kernels
Rubin uses the new GPU compile target sm107, while kernels built for Blackwell’s family target sm100f can run unchanged, according to vLLM’s support notes.
- Container:
vllm/vllm-openai:cu134-nightlyincludes CUDA 13.4 and PyTorch 2.15. - FlashInfer 0.7.0: Rubin-tuned attention, GEMM, and MoE kernels.
- MiniMax Sparse Attention: Rubin-tuned MSA prefill kernels are integrated.
- HBM4: about 2.4× GB200’s memory bandwidth.
- NVLink 6: about 1.7× GB200’s bidirectional bandwidth.
- Softmax exponentials: 2× FP32 and 4× BF16/FP16 throughput per SM per clock versus GB200.
- NVFP4 figures: the per-GPU comparison shows 3.5× dense training FLOPS, while its 5× inference comparison marks Rubin’s figure as sparse.
Locality-aware MoE
CUDA 13.4 exposes locality domains that place computation and data together, with SMs accessing nearby HBM at higher bandwidth and lower latency. vLLM describes its implementation in the MoE deep dive:
- Split FC1 and FC2 weights column-wise using a split-N strategy.
- Place each weight shard in its locality domain’s memory.
- Use Green Contexts and CUDA streams to launch one kernel on each domain’s SMs, restricting weight reads to the local shard.
- Keep the smaller decode activations non-localized across the domains.
Backfill mode, enabled with cudaDevSmResourceGroupBackfill, restores use of all 212 SMs. The preliminary FC1 + FC2 tests average roughly 1.2× faster execution in small-token settings, with balanced routing and communication time excluded.
SGLang and Miles
The SGLang and Miles teams tested two early-access Rubin nodes, eight GPUs, according to their separate report.
- Kimi K3 NVFP4 inference: MoE tail fusion improved end-to-end performance by 5.9% on Rubin.
- KDA speculative verification: the production verify kernel became 20% faster while preserving bitwise-identical outputs.
- Agentic RL: Miles used SGLang for rollouts and Megatron for training, with 64 concurrent sandboxes on a single tray’s Arm-based Vera CPU. The run served and trained Qwen3.5-35B-A3B on SWE-bench Verified tasks with arm64 images.