Skip to content

Query routing and model specialization: Nvidia strengthens the agent ecosystem with the Neuron 3.5 Lightning model and the Switchyard library

Share
Query routing and model specialization: Nvidia strengthens the agent ecosystem with the Neuron 3.5 Lightning model and the Switchyard library

Listen to this article

Read by Anchor

The global AI ecosystem is moving toward a shift from traditional conversational models to autonomous independent agent systems, which are systems that require continuous operation and reliance on an integrated network of models rather than a single primary model. In this context, Nvidia has expanded its open-source model family by unveiling the Neuron 3.5 Lightning model, a hybrid model that uses hybrid-expert technology and contains 30 billion parameters, designed specifically to deliver the highest levels of efficiency and speed for large-scale tasks within multi-model agent environments. This release builds on the previous Neuron 3 Nano series, aiming to give enterprises full control over where models run, how they are developed and adapted to their own data.

The Neuron 3.5 Lightning model is capable of delivering output up to four times faster than other models in its class, accelerating agent task completion by about 30 %.Benchmark evaluations using Bench-Bench tests show that the model can achieve accuracy comparable to leading large-scale models while reducing cost and time. The Neuron Alliance, which includes a range of institutions and technology companies, contributed evaluation methodologies, inference software and datasets that helped train and develop this fully customizable open model on Nvidia’s NeMo platform.

Leading technology companies have begun adapting the model for their specialized tasks, with CrowdStrike customizing it for cybersecurity work and security-alert monitoring, while Harvey, in partnership with Trajectory, deployed it for legal services, and CodeRabbit together with Bistane uses it for software review and code auditing. Lila Sciences has also adopted it to boost inference capabilities in physical and life sciences, and Fasteno Labs recorded advanced accuracy levels when using it for software development across financial and healthcare sectors. The model offers full flexibility in deployment and privacy options, as it can run locally on Nvidia RTX computers, DGX Spark platforms, DGX Station, Jetson, or be scaled out across workstations, data centers and cloud environments.

Alongside the new model, Nvidia launched the open-source NeMo Switchyard library for routing queries within smart agent tools.The library automatically analyzes each query and routes it to the most efficient and suitable model for the task, whether the model is open, proprietary, or a Nvidia model, without needing to rewrite applications. It allows developers to adjust routing algorithms to align with an organization’s priorities regarding quality, response time or cost reduction. Internal Nvidia evaluations revealed that routing queries through NeMo Switchyard preserves the same advanced accuracy while reducing task execution cost to roughly one-third of the cost of relying solely on the Opus 4.8 model.

Partner results from testing the NeMo Switchyard library showed tangible benefits in performance and actual spending. Boomi achieved 100 % routing accuracy across different domains while reducing response time in later rounds by 21 %, whereas Cadence improved operational efficiency by 9.9 % in formal verification use cases. Class Method recorded a 27 % cost reduction while maintaining quality, and Cognition integrated the directed model into a DevIn Desktop environment to achieve a 28 % average cost decrease. Meanwhile, LangChain cut costs by 74 % across 145 multi-round tasks by routing only 7 % of queries to large models, and RAMP reduced costs by 58 % and runtime by 33 % in programming evaluations.

To enhance transparency and auditability of open models, Nvidia announced the release of the model’s training data and techniques as permitted by the licenses, along with the Neuron RL-Agentic Terminal BeVOT reinforcement-learning dataset for training programming agents. The Neuron 3.5 Lightning model is currently available through the Hugging Face, ModelScope, OpenRouter platforms and as the precise NIM service on Nvidia’s website, while the NeMo Switchyard library is offered on GitHub to encourage adoption of multi-agent systems with higher operational and economic efficiency.

Don't miss the next story

Subscribe for updates