NVIDIA has released Nemotron 3.5 Lightning, an open 30-billion-parameter mixture-of-experts model with about 3 billion active parameters. It is built for the repetitive execution work inside long-running AI agents, such as calling tools, checking results and formatting outputs. Larger models can still handle planning and complex reasoning, while Lightning takes the high-volume steps that follow. NVIDIA reports the model produces output up to four times faster than similar-sized systems and completed 10,000 tasks 30 percent faster than Qwen3.6 35B on PinchBench at similar accuracy. The company also released NeMo Switchyard, an open-source library that routes each request to the most suitable model. Weights, training data and recipes are available under the OpenMDW-1.1 license. The model runs on hardware from Jetson and GeForce RTX systems to data centers and works with tools including Ollama, LM Studio and llama.cpp.
NVIDIA has released Nemotron 3.5 Lightning, an open 30-billion-parameter mixture-of-experts model with about 3 billion active parameters. It is built for the repetitive execution work inside long-running AI agents, such as calling tools, checking results and formatting outputs. Larger models can still handle planning and complex reasoning, while Lightning takes the high-volume steps that follow. NVIDIA reports the model produces output up to four times faster than similar-sized systems and completed 10,000 tasks 30 percent faster than Qwen3.6 35B on PinchBench at similar accuracy. The company also released NeMo Switchyard, an open-source library that routes each request to the most suitable model. Weights, training data and recipes are available under the OpenMDW-1.1 license. The model runs on hardware from Jetson and GeForce RTX systems to data centers and works with tools including Ollama, LM Studio and llama.cpp.