OZZZER · AI NEWS

Latest AI news.
Editorial picks.
AI News

Selected signals with a direct route back to each original source.

Share AI News
LinkedInXFacebookWhatsApp

Latest stories

Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each

Models & tools

Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each

Publisher preview

How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the answer: It uses a Mixture-of-Experts (MoE) architecture that selects only a subset of its parameters for each token. There are two dominant model architectures: Dense model and MoE. How a model organizes its parameters matters as much as how many it has. The right choice affects throughput, memory cost, and serving complexity more than raw parameter count does. Therefore, choosing between them comes down to…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
Translating CUDA Tile Operations from Python to Rust Using Agentic AI

Automation & Agents

Translating CUDA Tile Operations from Python to Rust Using Agentic AI

Publisher preview

cuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to tile-based GPU kernels, it splits mutable outputs into disjoint pieces and preserves the host-side ownership contract across kernel launches. It also allows programmers to opt out locally when they need lower-level control, enabling direct execution of Tile IR operations. The TileGym CUDA tile kernel library has accumulated a large library of production kernels written in CUDA Tile Python (cuTile Python) and Triton-TileIR (nvtriton). To make all…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

Research

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

Publisher preview

AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through a sequence of steps. It selects tools, evaluates their results, and continues reasoning within an increasingly long conversation. This workflow places new demands on edge inference. The model must generate tokens quickly, process long shared histories efficiently, and produce valid tool calls within a limited power and memory envelope. In the MLPerf Inference v6.1 Edge Agentic benchmark, NVIDIA TensorRT Edge-LLM ran Qwen3.6-27B on…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
How to Use AI Agents to Prepare 3D Scenes for Simulation

Automation & Agents

How to Use AI Agents to Prepare 3D Scenes for Simulation

Publisher preview

Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in OpenUSD, add physics properties, render preflight views, and validate the result against simulation-ready (SimReady) requirements. This workflow follows that process from a scene in Blender to a simulation-ready OpenUSD handoff for NVIDIA Isaac Sim or NVIDIA Isaac Lab. In practice, however, if you’re building agents for robotics, the workflow can often get stuck. It’s tempting to blame the hard part on the policy, the model,…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
Benchmarking LLM Inference at Scale with AIPerf

Coding & Development

Benchmarking LLM Inference at Scale with AIPerf

Publisher preview

You’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast? Your instincts might lead you to send curl commands, hand-roll an asyncio script, or vibe code yet another one-off load generator. All of these paths have the same problem: single-process performance limits, Python’s GIL capping concurrency, or numbers measured against a reference you built yourself. Either way, you end up with results you can’t fully trust, attached to tooling you’ll have to rewrite the moment requirements change. What you…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
How to Evaluate AI Agents From Tool Calls to Task Completion

Automation & Agents

How to Evaluate AI Agents From Tool Calls to Task Completion

Publisher preview

When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and recover when a step fails. Scoring whether the model sounds right tells you almost nothing about whether the work finished. That gap is why agent evaluation has had to evolve from scoring a single function call to scoring an entire task, with tool calling as the connective tissue underneath. This post traces that arc and explains why nearly every serious agent…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

Models & tools

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

Publisher preview

The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability that enables a single TensorRT network to execute across multiple GPUs using NCCL-backed distributed collectives while retaining TensorRT inference optimizations. It is fully supported starting with TensorRT 11.0. NVIDIA Dynamo-Triton (formerly NVIDIA Triton Inference Server) release 26.07 enables the multi-device inference capability of the TensorRT backend. One Triton KIND_MODEL instance can own multiple GPUs, create per-rank TensorRT execution contexts, CUDA streams, and NCCL communicators, and launch…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS

Automation & Agents

Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS

Publisher preview

GPU acceleration can speed up compute-intensive robotics workloads, but a fast CUDA kernel alone does not guarantee a fast ROS 2 graph. As messages move between nodes, they may continue to be serialized or copied through CPU memory, eroding the benefits of keeping perception and AI workloads on the GPU (Figure 1). With the upstream rosidl::Buffer abstraction and the CUDA buffer backend that NVIDIA recently contributed to ROS Lyrical, ROS 2 nodes can exchange GPU-resident payloads through zero-copy transport when runtime conditions allow, while preserving standard ROS 2 messages and…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing

Business use

Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing

Publisher preview

As large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated settings, data must be processed inside a trusted environment. NVIDIA Confidential Computing (CC) provides a pathway for running these workloads securely using memory-encrypted confidential virtual machines (CVMs), confidential GPUs, and encrypted NVIDIA NVLink. This enables running production AI inference on trusted hardware. Inference frameworks such as NVIDIA TensorRT LLM deliver best-in-class AI inference by combining framework-level optimizations with NVIDIA accelerated computing. However, when these frameworks run in a CC-enabled environment, secure…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
How SWE-Serve Exposes the Gap Between Local Tests and Live Serving

Coding & Development

How SWE-Serve Exposes the Gap Between Local Tests and Live Serving

Publisher preview

An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software therefore requires checking the full serving path, including whether the system returns correct results through its public interface. Developed with input from the SGLang team, SWE-Serve evaluates this gap with 53 tasks derived from merged changes to SGLang, an open-source system for serving large language models. Across 19 tasks with live-serving checks, the same patches passed 69.4% of the time when those checks were excluded,…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
Introducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought Reasoning

Healthcare

Introducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought Reasoning

Publisher preview

Radiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and data-dense modalities—the 3D computed tomography (CT) scan—remains largely underserved by modern vision language models (VLMs). Frontier general-purpose models perform poorly on volumetric imaging, and most open medical AI models lack the multistep conversational depth that radiologists need to trust and verify AI-generated findings. NVIDIA is addressing this gap with NV-Reason-CT, a VLM purpose-built for 3D CT analysis. NV-Reason-CT extends chain-of-thought reasoning to full volumetric CT,…

NVIDIA · 1 Oct 2026 · 20:31 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp
OpenAI cuts ties with 3 safety researchers, WSJ reports

Safety & Security

OpenAI cuts ties with 3 safety researchers, WSJ reports

Publisher preview

OpenAI has parted ways with three researchers on its safety team who allegedly shared confidential company information with a third-party AI safety organization, The Wall Street Journal reported on Thursday. “We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information,” an OpenAI spokesperson said in a statement to the WSJ. “Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work.” The report did not name the researchers, the…

TechCrunch AI · 1 Oct 2026 · 20:14 CESTRead story →Original ↗
Share
LinkedInXFacebookWhatsApp