OZZZER · AI NEWS1 of 3 free stories opened
← Back to AI News

Automation & Agents · 29 Sep 2026 · 23:06 CEST

Tracing Agent Harness Behavior with NVIDIA NeMo Relay

NVIDIA · 29 Sep 2026 · 23:06 CESTRead original at NVIDIA ↗
Share
LinkedInXFacebookWhatsApp
Tracing Agent Harness Behavior with NVIDIA NeMo Relay

Publisher preview · OZZZER analysis pending editorial review.

PUBLISHER ARTICLE PREVIEW

From the original article

Learn how to use execution traces to understand agent behavior and determine whether harness changes improve task outcomes.

An agent can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching the same content again. A correct final answer hides those extra steps, even though they increase latency and consume tokens. Inefficiencies create more chances for failure.

To improve an agent’s behavior, developers must understand whether a task succeeded and how the agent completed it. A success check by itself cannot explain why an agent recovered from a tool error, stopped early, or needed extra model calls.

In this tutorial, you’ll run two Hermes Agent examples with NVIDIA NeMo Relay. You’ll use the resulting traces to inspect model and tool calls, errors, retries, duration, and token use, then compare that evidence with each task’s verification result. A Hermes ToolPerf case study shows how to use the same approach to evaluate harness changes across repeated runs.

Next is an overview of the technologies used for this tutorial and how they work together.

NeMo Relay gives agent developers a common way to observe and control model and tool execution. The popular agent harness Hermes Agent includes NeMo Relay natively and represents its sessions, turns, model calls, and tool calls in NeMo Relay’s scope hierarchy. NeMo Relay records lifecycle events as work begins and ends, preserving its timing and parent-child relationships.

NeMo Relay is used for agent observability. You will work with three representations of agent execution:

An ATIF tool request shows what the model asked to run, but it does not confirm the outcome. To verify what happened, inspect ATOF for the matching tool start and end events and any recorded errors. Their shared uuid pairs the events, while parent_uuid connects the tool call to its parent.

Review traces before sharing them. Depending on your configuration, they can contain prompts, model responses, tool arguments and results, file paths, and other application data.

For agent safety and security governance, NeMo Relay helps provide the evidence layer: structured traces and trajectories that enterprises, evaluators, and security systems can use to investigate agent behavior, evaluate policies, improve controls, or create specialized security plugins that extend Relay.

The first example is intentionally small so you can verify the complete setup before adding web search and Phoenix. Hermes uses its terminal tool to run the included Python script inside an isolated Docker container. The script prints: VALUE=42.

That fixed output gives the runner an exact success check. A passing run also confirms that Hermes reached the model, invoked the terminal tool in the sandbox, and produced both Relay trace

Source

NVIDIA · 29 Sep 2026 · 23:06 CEST

Open the original at NVIDIA ↗