Frontier Flux Original

NVIDIA· 2 min read

NVIDIA ships Nemotron 3.5 Lightning: 30B MoE, 3B active, open weights

On 11 August 2026 NVIDIA released Nemotron 3.5 Lightning, an open 30B mixture-of-experts model with 3B active parameters for high-volume always-on agent execution under OpenMDW-1.1, plus NeMo Switchyard for multi-model routing.

ShareXEmail
Abstract cyan lightning arcs and sparse MoE node graph on charcoal

Abstract concept art for NVIDIA Nemotron 3.5 Lightning open agent model

Frontier Flux / AI-generated

NVIDIA shipped Nemotron 3.5 Lightning on 11 August 2026: open weights for a 30B mixture-of-experts model with 3B active parameters, NVFP4 and BF16 checkpoints on Hugging Face under OpenMDW-1.1, and NeMo Switchyard, an open library that routes agent steps across a mix of models.

From the 11 August NVIDIA AI post: "Introducing NVIDIA Nemotron 3.5 Lightning. An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models."

Key facts

  • Product surface: NVIDIA Developer Blog technical post and Hub card `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4` (11 August 2026). Corporate companion: Nemotron Lightning and NeMo Switchyard.
  • Architecture on the card: hybrid Mamba-2 + MoE + attention, up to 1M context, 30B total / 3B active. License OpenMDW-1.1; Hub marks the model ready for commercial use (full OpenMDW text still governs redistribute and product use).
  • NVIDIA reports up to 4x output speed vs similar-sized open models. On PinchBench, the developer post says Lightning reaches about 86% accuracy and completes tasks about 30% faster than Qwen3.6 35B at similar accuracy (NVIDIA harness). Card tables (MMLU Pro, GPQA, SWE-bench Verified, PinchBench, and others) are NVIDIA-measured under NeMo Gym / Evaluator recipes.
  • Same window: multi-token prediction and draft helpers, try paths on build.nvidia.com, OpenRouter, Ollama, and local DGX Spark / RTX recipes. Lightning is framed as the high-volume execution layer next to larger Nemotron 3 planners such as Nemotron 3 Ultra.

How to read

Use the X post as the alert and the developer blog plus Hub card as the ship surface. Speed and Pareto lines are NVIDIA-reported until independent benches land. Switchyard is a separate routing library shipped with the model window.

Independence

Frontier Flux is not affiliated with NVIDIA. We link official posts and the Hub card; we do not rank models or endorse vendor benchmarks.

ShareXEmail

Free

Get the morning brief

Free daily email of model releases, benchmarks, and lab notes from the last 24 hours across OpenAI, Anthropic, DeepMind, Meta, NVIDIA, and more. Capability-first, not company drama.

Opens Buttondown to finish signup. Confirm by email if asked. No ads. Unsubscribe anytime.

Back to latest · All Originals · Morning brief