Nvidia released Nemotron 3.5 Lightning on August 11, a 30 billion parameter open-weight model the company says runs up to four times faster than similarly sized models and cuts the cost of finishing an agentic task to roughly a third of what Anthropic’s Opus 4.8 costs on its own. Part of that claim now has independent backing. The rest still rests on Nvidia’s account alone.
Why it matters
This follows the same pitch Nvidia used in July, when it launched a nearly identical cost-and-speed campaign for Nemotron 3 Ultra with LangChain, claiming a 74% cost cut backed by named partners. Companies deciding on AI infrastructure are working from vendor-supplied efficiency numbers unless someone checks them first.
What an outside test actually found
Artificial Analysis, whose benchmark suite runs independently of Nvidia, measured Nemotron 3.5 Lightning at 24 on its Intelligence Index, nine points above the earlier Nemotron 3 Nano and roughly level with OpenAI’s gpt-oss-120b, a model about four times its size. On a pre-release endpoint served by DeepInfra, Artificial Analysis measured median output speeds near 670 tokens per second, faster than how similarly sized models are typically served today.
Introducing NVIDIA Nemotron 3.5 Lightning⚡
An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster.
It delivers up to 4x the output speed of similar-sized models. pic.twitter.com/ENWrZe76pU
— NVIDIA AI (@NVIDIAAI) August 11, 2026
Separately, Ramp Labs ran Nvidia’s new routing tool, NeMo Switchyard, against its own internal coding benchmark, Ramp SWE-Bench, and said routed agents matched single-model performance while cutting cost and runtime. Ramp’s own post didn’t put a number on that reduction. Nvidia’s blog credits Ramp with a specific 58% cost cut and 33% runtime cut, figures Ramp’s own account doesn’t repeat.
We tested NVIDIA NeMo Switchyard’s stage router for coding agents in Ramp SWE-Bench.
Routed agents showed comparable performance to single-model controls while substantially reducing costs and runtime. pic.twitter.com/AUo4Xy0AZF
— Ramp Labs (@RampLabs) August 11, 2026
What’s still only Nvidia’s word
The same blog post credits four other results to named partners: Boomi at 100% routing accuracy with 59% of its traffic sent to a fine-tuned model it says runs five times faster, Classmethod at a 27% cost reduction, LangChain at a 74% cost cut with a 6% accuracy tradeoff, and CodeRabbit and Harvey customizing the model for code review and legal work. None of the four have published their own version of those specific figures as of this writing. Harvey has posted publicly about training earlier Nemotron models, Super and Ultra, with the AI startup Trajectory, but nothing found ties that work to Lightning.
| Partner | Nvidia’s claim | Independently confirmed? |
|---|---|---|
| Artificial Analysis | Intelligence Index 24, ~670 tok/s | Yes, measured independently |
| Ramp Labs | 58% cost cut, 33% runtime cut | Direction confirmed; Ramp’s own post gives no percentage |
| Boomi | 100% routing accuracy, 59% traffic shift | Not published by Boomi |
| Classmethod | 27% cost reduction | Not published by Classmethod |
| LangChain | 74% cost cut, 6% accuracy tradeoff | Not published by LangChain for Lightning specifically |
| CodeRabbit | Customizing model for code review | Not published by CodeRabbit |
| Harvey | Customizing model for legal work | Harvey confirms work on earlier Nemotron models, not Lightning |
| Lila Sciences | Improved reasoning for science tasks | Confirmed as adopter; also an NVentures-funded company |
One of the six validators is also an investment
Lila Sciences, credited in Nvidia’s release with improving Lightning’s reasoning for scientific tasks, raised part of its $350 million Series A from Nvidia’s venture arm, NVentures, according to Lila’s own announcement. The release doesn’t mention that relationship.
Nvidia’s own documentation says more than its announcement did
The model card published alongside Lightning states that its benchmark scores were “measured by NVIDIA under a consistent harness” and “may differ from vendors’ self-reported numbers.” The same card describes Lightning’s architecture as a hybrid of Mamba-2, mixture-of-experts, and attention layers, more specific than the “MoE model” language Nvidia used in its own social media post. It also discloses that some synthetic training data, generated by other AI systems including OpenAI’s gpt-oss-120b, carried political and nationalistic slants that Nvidia says it filtered out, and that its training data skews toward male and White-coded references in the subset where demographic terms appear at all.
Nvidia has faced disclosure scrutiny over crypto exposure before. A federal shareholder class action alleging the company concealed more than $1 billion in GPU sales tied to cryptocurrency mining between 2017 and 2018 was certified to proceed and is moving toward trial, a separate case from anything in this release but part of the same pattern of what the company does and doesn’t say upfront.
A full third-party comparison against Lightning’s named competitors would settle most of what’s still open. So would Boomi, Classmethod, CodeRabbit, or Harvey confirming the specific figures attributed to them, and evidence that NeMo Switchyard, still listed as version 0.1.0 and alpha on GitHub, holds up at the production volumes Nvidia’s pitch assumes.
AI Disclosure: Cryip uses AI-assisted tools to help refine language — correcting spelling and grammar and simplifying complex terms for readability.
We do this to make crypto topics easier to understand for readers at all experience levels. AI does not draft facts, sources, or conclusions. Every article is reviewed and approved by a human editor before publication. Read our full AI Use & Content Policy.









