Skip to content
Original
Cline · Blog· Etisha Garg·· 08/11/2026SelectedAI score60

NVIDIA Nemotron 3.5 Lightning Is Now on Cline, Free to Use

Original title: NVIDIA Nemotron 3.5 Lightning just landed in Cline

AI overview

NVIDIA has released Nemotron 3.5 Lightning, a customizable open-source model built for persistent agents, and it's now available for free in Cline.

Why it matters

NVIDIA's new open-source model is free to use on Cline, so you can decide whether it's worth switching for high-frequency agent workloads.

Full text

NVIDIA Nemotron 3.5 Lightning just landed in Cline

NVIDIA just released Nemotron 3.5 Lightning, a new customizable open model built for always-on agents and it is now available in Cline for FREE.

Nemotron 3.5 Lightning is a 30B MOE model with 3B active parameters, distilled from NVIDIA’s frontier Nemotron 3 Ultra. It supports up to 1M token context window and is designed for always-on agents that need to move through a lot of model calls quickly. If you're using Cline, this matters.

One task, a lot of model calls

Some parts of an agentic coding task need deeper reasoning. Others are much more execution heavy like finding the right file, reading an implementation, running tests, checking an error, making the next edit.

But every one of those steps still calls a model.

NVIDIA's approach with Nemotron 3.5 Lightning is to make those repeated calls fast while maintaining the agentic capability needed for specialized tasks. The model is trained for popular agent harnesses and designed for high throughput tasks. .

That makes it an interesting fit for executing heavy Cline tasks where the agent may go through the read, edit, test, and retry loop many times before the work is done.

In Cline, speed compounds over the course of a task. One faster response is nice, but faster responses across dozens or hundreds of turns can change how quickly the whole agent loop moves.

A 30B model with 3B active parameters

Nemotron 3.5 Lightning uses a hybrid mixture-of-experts architecture with 30B total parameters and 3B active parameters during generation. It’s distilled from NVIDIA’s frontier Nemotron 3 Ultra model and supports a context window of up to 1M tokens.

NVIDIA's preliminary throughput numbers are where Lightning stands out.

Despite the smaller active footprint, Nemotron 3.5 Lightning scores 86.2 on PinchBench, 54.3 on SWE-Bench Verified, and 73 on AA-Omniscience non-hallucination. Compared to other leading open models of similar size, Nemotron 3.5 Lightning offers up to 4x higher throughput – placing it on the accuracy-speed Pareto frontier for high-volume agent workloads. 

NVIDIA Nemotron 3.5 Lightning just landed in Cline
Source: https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/

It is an open, customizable model that can be post trained for specialized workflows and deployed locally, at the edge, in the datacenter, or in the cloud.

Using Nemotron 3.5 Lightning with Cline

Prerequisites:

  • A Cline account (free to create)
  • Cline installed in your IDE, or the Cline CLI

Setup

  1. Open Cline
  2. Go to Settings
  3. Select Cline Usage-Billing as your API provider
  4. Select nvidia/nemotron-3.5-lightning from the FREE model dropdown
  5. Done!

Nemotron 3.5 Lightning is an interesting addition to the growing set of open models built around agentic workloads. Not every step in an agent loop needs the same kind of model. Lightning focuses on making the repeated calls inside that loop fast enough to keep the whole task moving.

Source: Cline · Blog · cline.ghost.io