AWS shows multi-agent evaluation with Bedrock AgentCore
AWS published a guide for building a Strands-based multi-agent supply chain decisioning system and evaluating it with Amazon Bedrock AgentCore Evaluations. The post covers built-in, custom, and explainability evaluators, arguing multi-agent systems must select correct tools, respect constraints, and explain decisions rather than only produce fluent responses.
Sources and evidence
Summary last validated Oct 5, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- agents
Anthropic ships Claude Code v2.1.296 with gateway code policies and subagent auto-compact
Anthropic released Claude Code v2.1.296. It adds a `code` key to the Claude apps gateway's managed policies, applying CLI settings in Claude Desktop's Code tab and enabling gateway mode. Subagent frontmatter and `--agents` gain `autoCompactWindow`, plus new env vars for workflow subagent models and overloaded-request retry delays. Several managed-hook and gateway sign-in bugs were fixed.
- models
AWS Adds Bedrock Models, Faster AgentCore Agents, Knowledge Base Sync
AWS's September 2026 recap covers Amazon Bedrock, Bedrock AgentCore, and Strands updates: broader model choice, faster serverless agents with built-in evaluation, and automated knowledge base syncing via native enterprise connectors. The post aggregates the month's releases for AI builders across AWS's model and agent tooling.
- agents
Postman runs Agent Mode for 40 million developers on Amazon Bedrock
Postman and AWS detailed the architecture behind Postman's Agent Mode, which serves 40 million developers on Amazon Bedrock. The post covers controlling tool sprawl, exposing schema-based reads, and treating context as the real bottleneck when running an AI agent at scale rather than in a demo.
- agents
NVIDIA Highlights Developers Using Frontier AI Agents With Omniverse Libraries
NVIDIA published a blog post describing how developers combine frontier AI models with NVIDIA Omniverse libraries to build simulation applications. The post says developers direct AI agents to assemble assets, connect physics and rendering, and verify scene behavior, supporting work such as exploring scenarios, investigating failures and improving designs.
- agents
Claude Code v2.1.295 adds hook failure blocking and gateway model controls
Anthropic released Claude Code v2.1.295, adding onFailure:"block" so hooks that fail to start, time out, or exit unexpectedly block the action. It also adds Program Status Protocol (OSC 7501) terminal support, optional per-upstream models lists with wildcards, upstream_ttfb_ms stream timeouts, and gateway login settings.