Your agent finishes the work, but slowly, and the obvious fix looks like a bigger model. Read its execution trace first: it shows whether the extra time rescues tasks the agent used to abandon or repeats finished work, and either answer can point to a narrower, cheaper change than a new model.
An agent trace is a timed, step-by-step record of every model and tool call in a run, with its results. Paired with an automated completion check, it shows how the agent reached its answer, not only whether it did. NVIDIA NeMo Relay records traces as an event log, a readable trajectory and OpenTelemetry spans.
The NeMo Relay tutorial sets the approval bar: a change must repeatably raise completion, or hold it while improving the reliability, latency or cost measure you targeted. FAIS operates existing agents this way: trace the current assistant on fixed cases, then fix the recurring cause before anyone budgets for a new model.
Can a slower agent be doing better work?
A slower agent can be finishing work it used to abandon. Nous Research's Hermes ToolPerf rerun on August 6, 2026 compared a pinned baseline with tool-layer fixes; Qwen3 Coder 30B shows the pattern:
| Qwen3 Coder 30B measure | Baseline | Tool fixes |
|---|---|---|
| Verified completion | 19/27 | 22/27 |
| Mean model calls | 3.8 | 4.9 |
| Mean duration | 27 s | 42 s |
Each of 9 tasks ran 3 times per model per arm (108 runs) under identical prompts, limits and checks. The arms differed by one batch of tool-layer changes (pinned configuration), so the result describes the batch, not each fix.
Those averages are one run's result. In an August 2 run of the same batch, with a different baseline, Qwen took 21% fewer turns with the fixes but completion fell from 74% to 67%; the maintainers note that the rerun does not simply reproduce it. With 3 repetitions per task, the per-task trace shows which extra effort bought completions and which bought nothing.
Claude Sonnet 4.5 completed 24 of 27 runs on baseline and 23 of 27 with the fixes: no real change, from a higher start than Qwen reached. Earlier August 2 testing found that the stronger model already recovered from most induced errors in one turn. Model choice still matters; the trace tells you whether it is the cause.
Which extra steps are worth keeping?
The Hermes task-level audit and scored task tables show four patterns:
- Keep recovery that earns completion. On the blocked-command task, Qwen's completion rose from 33% to 100% because recovery instructions let it continue after a parser block. As the NVIDIA and Nous Research authors put it, "recovering costs turns that giving up never spends."
- Trim rereading. On the large-file task, completion rose from 67% to 100%, but Qwen re-read more than it needed: keep the larger read, cut the repeats.
- Review tool messages that send the agent wandering. On case-insensitive search, Qwen's mean turns rose from 3.3 to 9.3 at unchanged completion; a zero-match probe message prompted extra searches in 2 of 3 repetitions.
- Separate tool and provider faults from model behaviour. Hidden-file search scored 0% for Sonnet on both arms and 33% then 0% for Qwen, a known tool gap. Provider-side chat-template breakage ended 5 baseline and 2 fixes Qwen runs, lowering its completion on both arms.
How do you choose the next change?
The NeMo Relay evaluation method gives the order:
- Fix a set of cases with an exact, automated completion check.
- Trace repeated baseline runs.
- Change one cause, such as a prompt, tool message or configuration, holding model, provider, inputs and limits constant.
- Run the same number of repetitions; compare verified completion first, then calls, time and cost.
What changes when the trace replaces the guesswork?
Consider an illustrative scenario: a Taiwan precision-parts maker whose order-intake assistant drafts replies to purchase orders, checking each line against the part master and current drawing revision in the ERP.
Before, the sales coordinator waits on every draft and the team prices a larger model. A trace of fixed purchase orders shows the assistant fetching the same drawing again for every line item that references it. The fix is narrow: keep the first read for that order and confirm drafts still cite the correct revision. The coordinator still approves every reply. Measure accepted drafts, plus time and cost per accepted draft, including failed attempts.
Some traces point past the tooling. When an agent keeps searching because a customer's part number, the ERP item code and the drawing revision look unrelated, the missing piece is a shared definition, not a larger model. FAIS's foundational business ontology maps customers, parts, revisions and orders and how your ERP and CRM records relate, so agents, through connectors built for each system, work from one definition instead of re-deriving it each run.
Start with one assistant and a frozen case set
Choose one assistant, agree its acceptance check with whoever signs off its output, freeze real cases and record repeated baseline traces. Treat traces as sensitive; depending on configuration, they can hold prompts, model responses, tool arguments, file paths and application data.
If your agent is slow and nobody can say why, that is a good reason to discuss your workflow with us.
Sources
- Tracing Agent Harness Behavior with NVIDIA NeMo Relay | NVIDIA Technical Blog
- Rerun 2026-08-06 — full reproducible battery, checked-in data
- Hermes ToolPerf August 6 rerun — scored task tables
- Tracking: core toolset performance batch — terminal & file-ops turn-efficiency (12 PRs) · Issue #77056 · NousResearch/hermes-agent · GitHub
Questions operators ask
An AI agent execution trace is a timed, step-by-step record of every model call and tool call in a run, with its results. It shows how the agent reached its answer, not only whether it did.
Does a slower AI agent mean the model is worse?
Not necessarily. In Nous Research's August 6, 2026 Hermes ToolPerf rerun, Qwen3 Coder 30B completed 22 of 27 runs with tool fixes versus 19 of 27, while mean duration rose from 27 to 42 seconds, largely by recovering tasks it used to abandon.
How do I test whether a change actually speeds up my agent?
Fix a set of cases with an automated completion check, trace repeated baseline runs, change one cause while holding model, provider, inputs and limits constant, then rerun the same repetitions. Compare verified completion first, then calls, time and cost.