
Amazon Bedrock prompt caching cuts input token costs up to 90%
Amazon Bedrock now supports prompt caching, which can reduce input token costs by up to 90% when the same context is repeatedly sent to foundation models. An AWS blog post details six scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integration.
Sources and evidence
Summary last validated Sep 15, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- agents
BlockRun and Incarna use Amazon Bedrock AgentCore payments for per-inference agent billing
Amazon Bedrock AgentCore payments lets AI agents pay for services on demand with infrastructure-enforced spending limits. Incarna's agents pay BlockRun for model inference one request at a time over x402, reducing the work of adding x402 payment support from months to days, according to an AWS Machine Learning Blog post.
- infrastructure
PyTorch details session-aware agentic inference with NVIDIA Dynamo
PyTorch published a blog post describing session-aware agentic inference with NVIDIA Dynamo. It explains that agentic workloads differ from single-turn chat: an agent session can include a large initial prefill, repeated model calls, and parallel subagents, changing the traffic an inference server handles.
- infrastructure
AWS details multi-team GPU sharing on SageMaker HyperPod
AWS published a reference architecture for sharing one Amazon SageMaker HyperPod EKS cluster across multiple teams. It uses AWS IAM Identity Center for authentication, per-team SageMaker Domains and Kubernetes namespaces for isolation, HyperPod Task Governance for fairness, and namespace-level cost allocation for chargeback.
- infrastructure
MIT's Christina Delimitrou targets data center energy efficiency
MIT Associate Professor Christina Delimitrou is rethinking how large cloud computing systems operate to reduce the environmental impact of data centers. Her work focuses on making these facilities more energy efficient as their environmental threat grows.
- models
Anthropic's Claude Haiku 5.5 launches on Amazon Bedrock and Claude Platform on AWS
Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. Anthropic says it is the fastest, most efficient model in the Claude 5.5 family, built for subagents and high-volume, cost-sensitive work, and costs roughly 75% less than Claude Haiku 4.5 for most tasks.