
NVIDIA Outlines Agent Evaluation Shift From Tool Calls to Task Completion
NVIDIA published guidance on evaluating AI agents, arguing that scoring whether a model sounds right reveals little about whether work finished. The post says agent evaluation must evolve from scoring a single function call to scoring an entire task, focusing on whether agents execute chains of sequential tool calls against live environments and recover when steps fail.
Sources and evidence
Attributed quotes
“When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and recover when a step fails.”
“Scoring whether the model sounds right tells you almost nothing about whether the work finished.”
“That gap is why agent evaluation has had to evolve from scoring a single function call to scoring an entire task”
Summary last validated Sep 28, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- agents
NVIDIA Highlights Developers Using Frontier AI Agents With Omniverse Libraries
NVIDIA published a blog post describing how developers combine frontier AI models with NVIDIA Omniverse libraries to build simulation applications. The post says developers direct AI agents to assemble assets, connect physics and rendering, and verify scene behavior, supporting work such as exploring scenarios, investigating failures and improving designs.
- open source
NVIDIA Details Five-Step Workflow for SimReady Robotics Assets
NVIDIA published a five-step workflow for converting CAD assets into SimReady robotics simulation assets using Omniverse libraries, SimReady Foundation specifications, and agentic NVIDIA skills. The process covers configuring and validating materials, collision geometry, joints, and physics properties beyond simple OpenUSD geometry conversion, preparing assets before robot behavior testing.
- agents
Claude Code v2.1.295 adds hook failure blocking and gateway model controls
Anthropic released Claude Code v2.1.295, adding onFailure:"block" so hooks that fail to start, time out, or exit unexpectedly block the action. It also adds Program Status Protocol (OSC 7501) terminal support, optional per-upstream models lists with wildcards, upstream_ttfb_ms stream timeouts, and gateway login settings.
- agents
BlockRun and Incarna use Amazon Bedrock AgentCore payments for per-inference agent billing
Amazon Bedrock AgentCore payments lets AI agents pay for services on demand with infrastructure-enforced spending limits. Incarna's agents pay BlockRun for model inference one request at a time over x402, reducing the work of adding x402 payment support from months to days, according to an AWS Machine Learning Blog post.
- agents
NVIDIA KGMON Team Places Second in KDD Cup 2026 Data Agents Competition
NVIDIA's KGMON team placed second in the KDD Cup 2026 Data Agents competition. The team built a system around making an agent's harness smaller, clearer, and easier to verify. The competition required agents to answer natural-language questions across heterogeneous sources including databases, CSV and JSON files, prose documents, PDFs, and briefing videos.