Articles from Source: Nvidia-Developer-Blog

Evaluating AI Agent Skill Performance with NVIDIA SkillEvaluator

2026-08-19 16:00
NVIDIA has introduced the SkillEvaluator, an open-source tool designed to assess AI agent performance through skill evaluations. It enables comparison of agent output with and without specific skills, enhancing efficiency in task execution. Over 300 verified skills from more than 30 NVIDIA products have been benchmarked. Skills undergo a three-tier evaluation process, focusing on safety, distinctiveness, and live performance testing to ensure readiness. Discover more about AI performance...
Source: Nvidia Developer Blog
Michelle Horton

Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control

2026-08-19 16:00
🚀 Exciting advancements in robot control with the NVIDIA Cosmos 3 Edge! This new model adapts to sensors and environments while running on onboard hardware, providing solutions for on-device deployment. Key features include: - 4B omni-model with a 2B reasoner - Pretrained on physical-world data You'll learn to post-train this model for robot actions, run it on Jetson Thor, and evaluate its performance in simulation. Check out the open cosmos-framework repo for more details! #Robotics #AI...
Source: Nvidia Developer Blog
Michelle Horton

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit

2026-08-18 18:00
Unlock the potential of materials simulation with the NVIDIA ALCHEMI Toolkit! This innovative tool addresses challenges in atomistic simulations by combining machine learning with GPU acceleration. It streamlines the creation of simulation workflows using user-friendly, natural-language prompts. Explore how to set up your environment and improve code reliability with the Toolkit for effective simulation results. #NVIDIA #ALCHEMI #MaterialsScience #AICoding #SimulationTools 🧪💻🔬✨
Source: Nvidia Developer Blog
Elizabeth Goodman

Run Massive-Scale UMAP in Minutes Using Multiple GPUs—Without Losing Accuracy

2026-08-18 16:48
Unlock the power of UMAP with NVIDIA's latest advancements! 🚀 Uniform Manifold Approximation and Projection (UMAP) is essential for data visualization and feature extraction. However, as datasets grow, running UMAP can become costly and time-consuming. NVIDIA's cuML now allows multiple GPUs to speed up the all-neighbors graph construction, drastically reducing training time. This enables analysis of massive datasets in minutes rather than hours. Explore how to leverage this multi-GPU...
Source: Nvidia Developer Blog
Tanya Lenz

Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer

2026-08-17 18:12
Unlocking new possibilities with the Nemotron 3.5 Lightning NVFP4! ⚡ Developers are customizing models to meet specific targets for latency, speed, and memory. The latest checkpoint allows for up to 4x faster throughput while maintaining accuracy. This is achieved through quantization-aware distillation (QAD), which enhances model performance even with aggressive quantization methods. Explore how QAD improves memory efficiency and accuracy in the NVIDIA Model Optimizer. #NVIDIA...
Source: Nvidia Developer Blog
Tanya Lenz

Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

2026-08-12 18:23
🚀 Alibaba has unveiled Qwen3.8-2.4T-A95B, its largest open-weight model, featuring 2.4 trillion parameters and 95 billion activated per token. This model supports complex tasks like coding and document analysis with a unique mixture of full and linear attention for enhanced efficiency. NVIDIA collaborates to optimize its deployment for high-performance computing, achieving over 4K tokens per second per GPU. #AI #MachineLearning #Alibaba #NVIDIA #OpenSource
Source: Nvidia Developer Blog
Michelle Horton

How to Choose Full-Stack Observability for NVIDIA AI Factories

2026-08-12 16:13
Unlock the potential of NVIDIA AI factories with full-stack observability! 🌐 This article outlines a framework to connect multiple infrastructure layers, enabling teams to detect and isolate issues effectively. It highlights a case study on identifying gray failures in distributed training jobs due to hardware degradation. Learn how to map components to telemetry tools, prioritize alerts, and streamline monitoring with a single triage dashboard. Understanding these processes is crucial for...
Source: Nvidia Developer Blog
Jorge Cardoso

NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation

2026-08-11 19:00
🚀 NVIDIA has released JetPack 7.2.1, enhancing video capabilities across Jetson applications. This update introduces support for PyNvVideoCodec 2.2, allowing developers to leverage hardware-accelerated video encoding and decoding with AI-friendly features. The new agentic video skills streamline workflows, making it easier to connect goals with live device discovery and configuration. #NVIDIA #JetPack #VideoTech #AI #Jetson
Source: Nvidia Developer Blog
Elizabeth Goodman

NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents

2026-08-11 13:01
🚀 Exciting news in AI! NVIDIA has introduced the Nemotron 3.5 Lightning, a 30B mixture-of-experts model designed for high-volume execution in long-running AI agents. Its smaller design focuses on fast, low-latency task execution. This model enhances applications by optimizing tool calls, result validation, and subagent delegation. With innovative routing through NVIDIA NeMo Switchyard, tasks are intelligently assigned to the best models. #AI #NVIDIA #Nemotron #MachineLearning #TechInnovation
Source: Nvidia Developer Blog
Tanya Lenz

Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard

2026-08-11 13:00
🚀 Building AI agents involves more than just selecting a single model. Each model has unique strengths and weaknesses, impacting cost and performance. 🔄 NVIDIA NeMo Switchyard addresses this by routing tasks to the most suitable models, improving efficiency and accuracy without needing to redesign applications. 📊 This system allows developers to create more effective AI workflows, adapting to various workload requirements. #AI #NVIDIA #NeMoSwitchyard #MachineLearning #Efficiency
Source: Nvidia Developer Blog
Michelle Horton

Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA

2026-08-10 13:27
🚀 Meta has launched Muse Glimmer, a 30B open-weight model designed for local AI workflows. This model features a 120K+ context window and is optimized for various NVIDIA platforms, achieving 20K tokens/sec on a single GPU. Unlike typical chat-focused models, Muse Glimmer supports complex, multi-step tasks with high reliability and consistency. #Meta #AIModels #OpenSource #NVIDIA #TechInnovation
Source: Nvidia Developer Blog
Michelle Horton

Beyond VLAs: How World Action Models Reshape Robot Manipulation

2026-08-04 16:00
🚀 A recent article discusses advancements in robotics, focusing on the shift from vision-language-action (VLA) models to world action models (WAM). WAMs address the challenge of generalizing policies by incorporating a video world model, enhancing a robot's ability to predict scene dynamics, unlike VLAs that primarily focus on semantic understanding. This evolution allows robots to better adapt to new conditions and tasks by leveraging physics knowledge. The NVIDIA Cosmos 3 model serves as a...
Source: Nvidia Developer Blog
Michelle Horton

Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super

2026-08-04 15:00
🚗 NVIDIA introduces Alpamayo 2 Super, a powerful 34-billion-parameter model for autonomous vehicle (AV) development. This technology integrates trajectory generation, intent prediction, scene understanding, and data labeling into one system, streamlining the development process. Alpamayo 2 Super offers 360-degree perception with multiple outputs including future trajectories and reasoning auto-labels. Developers can utilize the same model across various tasks, enhancing efficiency. Check out...
Source: Nvidia Developer Blog
Elizabeth Goodman

How to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure

2026-08-03 16:00
🚀 Running isolated tenant Kubernetes clusters on shared GPU infrastructure can enhance team autonomy without unnecessary hardware splits. This approach utilizes a single control plane with GPU sharing and per-team quotas. By using KAI Scheduler and vCluster, teams can operate independently while sharing resources effectively. The tutorial demonstrates how to set up three teams with their own GPU Kubernetes pods on one physical GPU, ensuring workload visibility is limited to each team. Learn...
Source: Nvidia Developer Blog
Tanya Lenz

NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage

2026-08-03 16:00
NVIDIA's new Vera BlueField-4 STX Storage Processor enhances AI-native storage systems. It improves encryption, compression, integrity checking, and recovery, addressing the growing demands of concurrent AI agents. Faster processing reduces CPU load and power use while increasing efficiency. The benchmark shows Vera outperforming traditional x86 CPUs, optimizing data flow and performance. #NVIDIA #AI #DataStorage #TechNews #Innovation 🚀💻📊🔒
Source: Nvidia Developer Blog
Elizabeth Goodman

Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference

2026-07-31 22:16
Exploring AI model design is crucial as long-context workloads rise. 🧠 This article highlights the significance of attention in inference performance. It discusses how model architecture should align with GPU execution for optimal results. Key factors like group size, head dimension, and sequence length are examined, leading to practical guidelines for enhancing throughput and interactivity on NVIDIA GPUs. 📈 Stay tuned for insights on sparse attention! 🔍 #AI #MachineLearning #ModelDesign...
Source: Nvidia Developer Blog
Tanya Lenz

NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek

2026-07-31 15:13
🚀 Exciting news for developers! NVIDIA Video Codec SDK 13.1 is now available, designed to enhance video pipelines for high-quality streaming and content delivery. Key updates include: - **New encode features**: Hierarchical Reference Mode for AV1 with up to 31 B-frames. - **Decode enhancements**: Per-macroblock decode stats for H.264 and HEVC. - **Transcode improvements**: Redesigned modular transcoder samples. Explore these features and provide feedback on the NVIDIA Developer forums! 💻🎥...
Source: Nvidia Developer Blog
Elizabeth Goodman

Run High-Performance Core Math at Scale with NVIDIA nvmath-python

2026-07-30 22:43
Unlock high-performance math in Python with NVIDIA's nvmath-python! 🚀 This new library connects the Python scientific community with NVIDIA's powerful CUDA-X math libraries, allowing seamless access to optimized math operations on CPUs and GPUs. Version 1.0 offers easy installation and customization options, making it simple to integrate into your workflows. nvmath-python enhances existing array libraries like NumPy and PyTorch, focusing on GPU-accelerated routines without deep coding...
Source: Nvidia Developer Blog
Michelle Horton

Four Ways to Deploy More Secure AI Agents

2026-07-30 21:09
🔍 Knowledge workers are increasingly using AI agents as digital coworkers to boost productivity. These agents can manage tasks like bug fixes and testing, but they also present security risks. Recent evaluations by NVIDIA's AI Red Team identified common vulnerabilities, including lack of access control and exposure of sensitive information. To enhance security, it’s crucial to implement strong access controls and limit code execution capabilities. Understanding these risks can help...
Source: Nvidia Developer Blog
Michelle Horton

NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure

2026-07-30 16:00
Unlocking AI Infrastructure Performance! 🚀 NVIDIA's latest insights reveal that two identical AI clusters can show performance gaps of 8% to 12% due to configuration choices. These include settings in the kernel, hypervisor, and NVIDIA NCCL, which can impact training throughput. The article discusses four diagnostic investigations that pinpoint issues in system memory management, power management, and more. Infrastructure engineers can use these patterns to enhance their own clusters before...
Source: Nvidia Developer Blog
Elizabeth Goodman

How to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo Guardrails

2026-07-29 16:46
🚀 Deploying an AI coding assistant in regulated environments presents key challenges, including network restrictions and supply-chain risks. This tutorial outlines how to self-host a validated coding assistant using NVIDIA infrastructure. You'll learn to set up a StarCoder2-7B NIM endpoint and implement NVIDIA NeMo Guardrails for policy enforcement. Key prerequisites include an NGC API key and a compatible NVIDIA GPU. The tutorial emphasizes a modular design, allowing teams to adopt...
Source: Nvidia Developer Blog
Tanya Lenz

Developing Healthcare Robotics with GPU-Native Medical Physics Simulation

2026-07-28 20:49
🚀 Developing healthcare robotics presents unique challenges that differ from other fields. Firstly, there's a significant data gap. Most teams lack the vast datasets needed for training, particularly for rare cases that impact clinical safety. Secondly, generalization is limited. While imitation learning has its place, reinforcement learning offers a way to explore more scenarios through realistic simulations. Lastly, development velocity is slow, often taking several years due to the...
Source: Nvidia Developer Blog
Michelle Horton

NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning

2026-07-27 16:00
🚀 NVIDIA has released Ising Calibration 1.5, an open-source vision language model aimed at automating quantum computer calibration. This model interprets diagnostic outputs from quantum processors and adapts to unfamiliar data without prior examples. It is 11.4% smaller and can be deployed on a single GPU or NVIDIA DGX Spark. Ising Calibration 1.5 shows strong performance on the QCalEval benchmark, significantly improving over its predecessor and remaining competitive with leading models....
Source: Nvidia Developer Blog
Tanya Lenz

Six Agent Harness Capabilities for Higher Model Performance

2026-07-27 09:00
Building effective AI agents extends beyond model selection; it involves the architecture around the model. The design of this harness can significantly influence performance outcomes. 🛠️ NVIDIA Labs has introduced the NOOA framework, an open-source tool that simplifies agent development using a single Python class for easier coordination of capabilities and state management. 🐍 This approach allows for reproducible research and community collaboration. Check out the potential of NOOA in...
Source: Nvidia Developer Blog
Michelle Horton

Advancing Semiconductor Innovation Across Materials Engineering and Manufacturing

2026-07-27 00:45
The semiconductor industry faces rising demands due to increased AI workloads. 🖥️ Meeting performance targets is crucial, as even minor delays can lead to significant financial impacts. To address these challenges, Applied Materials and NVIDIA are collaborating on a digital development model. This partnership enhances materials engineering and semiconductor manufacturing through advanced simulations and AI-driven insights. Their approach includes GPU-accelerated simulations, physics-based...
Source: Nvidia Developer Blog
Tanya Lenz

NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding

2026-07-27 00:45
NVIDIA's Nemotron 3 Ultra and the ACE-RTL agent are transforming RTL coding efficiency and accuracy. 🖥️ As chip design faces time constraints, these tools enhance code generation and error correction through iterative testing and feedback. The CVDP benchmark offers a realistic assessment of LLMs in RTL tasks, focusing on complex coding scenarios. 🔄 With agentic workflows, engineers can effectively tackle RTL challenges by reusing and modifying code, interpreting failures, and debugging. 🔍...
Source: Nvidia Developer Blog
Elizabeth Goodman

ModelExpress: Distributing Model Artifacts at the Speed of Light

2026-07-24 16:45
🚀 ModelExpress is revolutionizing how we distribute model artifacts. As model sizes grow, transferring weights can be costly and time-consuming. ModelExpress addresses this by optimizing the transfer path, utilizing existing weights in GPU memory to speed up the process. By transferring weights directly from GPU to GPU, MX significantly reduces startup times—from 8 minutes to just 1 minute 44 seconds! This efficiency extends to kernel caches and RL weight updates as well. #ModelExpress #AI...
Source: Nvidia Developer Blog
Elizabeth Goodman

Debugging Ray Tracing Applications Using NVIDIA OptiX Toolkit

2026-07-23 16:07
🔍 Debugging ray tracing applications can be challenging, especially with the NVIDIA OptiX toolkit. The OptiX ray tracing engine offers tools to diagnose issues like invalid API arguments or GPU-side bugs. The NVIDIA OptiX Toolkit (OTK) is a GitHub resource that helps developers streamline their debugging process. Key features include consistent checking of OptiX and CUDA API error codes and targeted device-side debug printing. OTK also provides an example program, DemandPbrtScene, to...
Source: Nvidia Developer Blog
Tanya Lenz

Start Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes

2026-07-23 16:00
Unlock the power of customization with NVIDIA Nemotron 3 Nano! 🖥️ This article highlights how developers can tailor models to specific use cases, despite challenges like infrastructure and expertise. The NVIDIA Nemotron 3 family, paired with Prime Intellect Lab, simplifies this process. In just five minutes, you can set up a customized model using hosted reinforcement learning. The tutorial guides you through the process with a simple Python Math example. #NVIDIA #AI #Customization...
Source: Nvidia Developer Blog
Michelle Horton

Make Long-Running NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++

2026-07-22 16:35
Developers using NVIDIA TensorRT may face challenges during engine builds, which can take from seconds to several minutes. Long builds, especially with large models and new GPU SKUs, can leave users unsure of the process status. Current integrations often do not report progress or allow for cancellation, leading to wasted GPU resources. Improving observability and cancelability in TensorRT builds could enhance workflow efficiency. 🖥️⏳ #NVIDIA #TensorRT #AI #MachineLearning #GPU
Source: Nvidia Developer Blog
Michelle Horton

Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI

2026-07-21 15:00
🚀 The NVIDIA Rubin GPU architecture is transforming AI capabilities by enabling continuous, large-scale intelligence production. These advancements support complex tasks with sustained reasoning, requiring efficient processing and low latency. The Rubin platform features enhanced Tensor Cores and a powerful memory subsystem, achieving significant improvements in energy efficiency and performance. #NVIDIA #AI #TechInnovation #GPUs #AgenticAI
Source: Nvidia Developer Blog
Tanya Lenz

NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI

2026-07-21 15:00
NVIDIA’s new Vera CPU features the Olympus core, designed for optimal single-thread performance in Agentic AI. This CPU shifts critical execution tasks to enhance responsiveness and throughput in AI operations. It focuses on strong single-thread performance, memory bandwidth, and predictable latency. Olympus was co-designed with the entire Vera Rubin platform to maximize efficiency across AI infrastructure workloads. For more technical insights, check out the NVIDIA Vera CPU white paper....
Source: Nvidia Developer Blog
Michelle Horton

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

2026-07-21 15:00
NVIDIA has set a new world record for pre-training with its GB300 NVL72, achieving 1,648 TFLOPs per GPU while training the DeepSeek-V3 model, which has 671B parameters. 💻 This advancement highlights the shift to mixture of experts (MoE) architectures, which allows for more efficient computation by activating only a subset of parameters per token. 🔍 However, the need for extensive communication between GPUs poses challenges, as delays can affect throughput during training. 📈 #NVIDIA #AI...
Source: Nvidia Developer Blog
Kirthi Devleker

NVIDIA NVLink: The Scale-Up Network for AI Factories

2026-07-20 15:46
NVIDIA NVLink is revolutionizing AI factories by addressing the increasing demand for complex AI workloads. As models grow larger, the need for high-bandwidth, low-latency GPU communication becomes essential. NVLink facilitates this, allowing multiple accelerators to operate as a cohesive unit. This technology is crucial for efficient AI inference and training, ensuring resilience and high performance in data centers. #NVIDIA #AI #DataCenter #TechInnovation #NVLink 🤖💻🔗
Source: Nvidia Developer Blog
Elizabeth Goodman

Integrate NVIDIA Omniverse RTX Sensor Simulation Into Existing Apps

2026-07-20 15:00
Developers in 3D design, simulation, and robotics can enhance their applications with NVIDIA Omniverse RTX Sensor Simulation. The new ovrtx library allows for generating sensor outputs from OpenUSD content using a lightweight SDK. This integration supports seamless workflows, enabling teams to visualize data and validate systems effectively. Learn more about optimizing your applications with NVIDIA's tools! 🌐🔧 #NVIDIA #Omniverse #3DDesign #Simulation #AI
Source: Nvidia Developer Blog
Tanya Lenz

Q&A: How Capcom Brought Path Tracing to RE ENGINE Across PRAGMATA and Resident Evil Requiem

2026-07-16 22:59
Capcom's RE ENGINE team has successfully integrated path tracing into Resident Evil Requiem and PRAGMATA, enhancing their visual experiences. Over two years, they developed a game-oriented path tracer optimized for DLSS, improving direct lighting and reducing visual gaps between gameplay and cutscenes. 🌟 This transition reflects a commitment to immersive gameplay, leveraging advanced NVIDIA technologies for realistic lighting and reflections. #Capcom #REENGINE #PathTracing #Gaming #ResidentEvil
Source: Nvidia Developer Blog
Michelle Horton

Integrating Context-Aware Video AI Agents Into Enterprise Workflows

2026-07-16 16:03
Unlocking the potential of video analytics in enterprises involves integrating AI agents into existing workflows. 🤖 The article discusses the challenges of merging video systems with tools like content management and messaging platforms. It highlights NVIDIA NemoClaw, which supports context-aware analysis and enables actionable insights from video data. Key features include generating structured reports and building multi-step workflows that enhance business processes. Learn about NVIDIA...
Source: Nvidia Developer Blog
Tanya Lenz

Scaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueField

2026-07-16 16:00
🚀 Agentic AI transforms how AI factories operate, enabling one request to trigger multiple model calls and data processes. NVIDIA's BlueField platform plays a crucial role by offloading tasks from CPUs, enhancing data movement, and ensuring policy enforcement. This leads to improved GPU utilization, reduced latency, and cost efficiencies. The combination of BlueField-4 DPUs and NVIDIA DOCA software supports advanced infrastructure needs, making data management integral to AI inference. #AI...
Source: Nvidia Developer Blog
Michelle Horton

Build a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 Skills

2026-07-15 23:00
🚀 Exciting advancements in video analytics are here with NVIDIA DeepStream 9.1! Developers can now track objects seamlessly across multiple camera views, thanks to the new Multi-View 3D Tracking (MV3DT) and AutoMagicCalib (AMC) features. These innovations automate camera calibration and provide consistent object ID tracking. DeepStream 9.1 supports 13 new agentic skills, enhancing the development of vision AI applications. With open-source access on GitHub, getting started has never been...
Source: Nvidia Developer Blog
Elizabeth Goodman

Develop Lightweight USD Runtimes Faster with AI Agents

2026-07-15 21:57
Unlock the potential of lightweight USD runtimes with AI agents! 🌐 OpenUSD is a framework that allows teams to integrate CAD data, simulations, and real-world data into a unified view. Traditionally, creating a USD implementation required extensive code adaptation. Now, with nanousd-labs, developers can generate runtimes directly from the USD Core Specification. This method streamlines the implementation process, making it faster and more efficient for physical AI applications. Explore how AI...
Source: Nvidia Developer Blog
Michelle Horton