IT Brief Canada - Technology news for CIOs & IT decision-makers
Canada
NVIDIA unveils Nemotron 3.5 Lightning for local agents

NVIDIA unveils Nemotron 3.5 Lightning for local agents

Wed, 12th Aug 2026 (Today)
Mark Tarre
MARK TARRE News Chief

NVIDIA has expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, a new open model aimed at local AI agents.

The launch is part of a broader push around open models, local inference tools and software for developers running AI workloads on PCs, workstations and compact DGX systems. Alongside the model release, NVIDIA outlined updates to clustering software, model support and local agent tooling.

Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model with open weights, allowing developers to fine-tune it for specific uses. It is intended for "always-on" agents that handle repeated or specialised tasks on local systems rather than relying solely on remote cloud services.

The model runs across a range of NVIDIA hardware, including RTX PCs, DGX Spark, Jetson devices, workstations and larger systems. Partners including Acer, ASUS, Dell, Exxact, GIGABYTE, HP, Lenovo, MSI and Supermicro are supporting Blackwell-based systems that can run the software.

Model push

The announcement also included a broader list of open-source and open-weight models now available or newly supported on NVIDIA hardware. They include Cosmos 3 Edge, a 4-billion-parameter world model for robotics and vision applications; MiniMax-H3, a 33-billion-parameter model that generates video and stereo audio; and Poolside AI's Laguna S 2.1 coding model.

NVIDIA also highlighted refreshed support for DeepSeek-V4-Flash and Thinking Machines Lab's Inkling-Small, both large models designed to run with reduced active parameter counts. In each case, optimised checkpoints or community-built formats are intended to make local deployment possible on systems such as DGX Spark and DGX Station.

Another part of the rollout focuses on agent software. Developers can build agents around local models with tools including NemoClaw, and fine-tune some models with NeMo Automodel while keeping data on the device.

Meta's Muse Glimmer also featured in the update. The 30-billion-parameter open-weight model is designed for coding and local agent tasks, and can run on a single GeForce RTX 5090 while handling long context windows and multistep workflows.

Local systems

A central part of NVIDIA's message is that larger open models can increasingly run on desktop or compact workstation hardware. To support that, NVIDIA updated its NVIDIA Sync software, which lets users cluster multiple DGX Spark systems together.

According to NVIDIA, the Sync application is available for Windows and macOS and can automatically detect connected systems, provide remote access and launch applications across one or more DGX Spark units. A Cluster Assistant within the software configures two or more systems as a high-speed cluster through ConnectX-7 ports.

NVIDIA said the updates are intended to help developers run larger models such as GLM 5.2 and DeepSeek V4 Flash when a single system does not provide enough memory or throughput. It also disclosed two additional DGX Spark software updates: a native ARM64 Linux build of Google Chrome and a Sync Resource Monitor for system-level usage tracking.

Video tools

NVIDIA also highlighted third-party creative software tuned for its chips. LTX 2.5, an open-world video generation model from LTX, has been optimised for RTX GPUs, DGX Spark and DGX Station systems. NVIDIA said it delivers up to 20% faster performance and 40% memory savings on an RTX 6000 PRO GPU.

Alibaba's Wan-Animate-2, a 14-billion-parameter model that transfers motion and facial expressions from video to a static character image, now has day-zero support in ComfyUI. NVIDIA said the model runs up to 16 times faster on an RTX PRO 5000 Blackwell and 26 times faster on an RTX 5090 than on Apple's M3 Ultra.

Unsloth Desktop is also launching as a fully open-source desktop application for local model inference, training, diffusion workloads, agent integrations, web research and code execution. NVIDIA said the software combines local AI training and inference in one desktop app.

Cost and routing

Alongside Nemotron 3.5 Lightning, NVIDIA introduced NeMo Switchyard, an open-source routing library for agent workflows. The software is designed to direct each stage of a task to different models based on speed, accuracy and cost.

NVIDIA said this approach addresses a growing concern among businesses deploying generative AI: token costs when every task is sent to a top-tier model. Internal benchmarks, according to the company, showed that routing workloads across several models with Switchyard reduced benchmark completion cost to roughly one-third of using Opus 4.8 alone while maintaining frontier-level task completion.

NVIDIA said it has worked with vLLM, Ollama, llama.cpp and LM Studio to support local deployment options for Nemotron 3.5 Lightning in NVFP4 and GGUF formats. Unsloth is also providing support through Unsloth Studio.

NVIDIA framed the broader set of updates as evidence that open-source software and open-weight models are making local AI more practical for developers and enthusiasts. The effort now spans coding agents, robotics, media generation and multimodal systems across hardware ranging from consumer GPUs to clustered desktop appliances.