
NVIDIA details NVLink 6 multi-layer resiliency for AI factories
NVIDIA published a developer blog explaining how NVLink 6 provides multi-layer resiliency for large-scale AI factories. The post notes that in massive AI training, every GPU must synchronize gradients across thousands of collective operations per second, while during inference unplanned downtime directly reduces requests served and limits revenue generation.
Sources and evidence
Summary last validated Sep 15, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- infrastructure
Ai2 replaces priority GPU scheduler with budget-based fair-share system
Ai2's AI Infrastructure team replaced its priority-based GPU scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and time-slicing contracts. The institute manages thousands of NVIDIA H100, B200, and B300 GPUs across 88- to 1024-GPU clusters for about 150 researchers, with demand running 2-3x available capacity. The change moved GPU allocation debates into a transparent administrative budgeting process.
- infrastructure
Ai2 details GPU scheduler using time budgets and fair-share allocation
Ai2 published a blog post explaining its new GPU scheduler for its clusters. The scheduler combines time budgets, fair-share allocation, and time-slicing to prioritize high-impact research, shorten queue waits, and keep GPUs busy.
- infrastructure
Cohere Publishes Guidance on Shared vs Dedicated Inference for Embed and Rerank
Cohere published a blog post explaining how to choose between shared, consumption-based inference and dedicated, provisioned inference for Embed and Rerank workloads. It argues request shape matters more than request volume, that traffic patterns determine cost efficiency, and that break-even depends on request size, token volume, and reranking candidate count rather than infrastructure pricing alone.
- agents
NVIDIA Highlights Developers Using Frontier AI Agents With Omniverse Libraries
NVIDIA published a blog post describing how developers combine frontier AI models with NVIDIA Omniverse libraries to build simulation applications. The post says developers direct AI agents to assemble assets, connect physics and rendering, and verify scene behavior, supporting work such as exploring scenarios, investigating failures and improving designs.
- open source
NVIDIA Details Five-Step Workflow for SimReady Robotics Assets
NVIDIA published a five-step workflow for converting CAD assets into SimReady robotics simulation assets using Omniverse libraries, SimReady Foundation specifications, and agentic NVIDIA skills. The process covers configuring and validating materials, collision geometry, joints, and physics properties beyond simple OpenUSD geometry conversion, preparing assets before robot behavior testing.