MLPerf Introduces End-to-End RAG Inference Benchmark
MLCommons announced the MLPerf End-to-End RAG Inference Benchmark, covering the full pipeline from vector database construction to serving iterative, multi-hop question-answering inference. This benchmark aims to standardize performance evaluation for retrieval-augmented generation systems.
Sources and evidence
Summary last validated Aug 29, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- industry
MLCommons publishes business guide to measuring and managing AI risk
MLCommons released a guide titled "Assessing & Managing AI Risk for Business and Commercial Deployments," aimed at business owners. It argues teams managing AI risk face tension between protecting the business and adopting AI faster, and that independent measurement can reconcile the two. The guide frames AI risk as a cost that can be measured, priced, and managed.
- research
NIST Study Maps How Toxic Adulterants in Fentanyl Vary Across US
NIST researchers analyzed fentanyl samples nationwide and found that the toxic substances mixed into the drug vary by region and shift over time. The study aims to help first responders and law enforcement identify dangerous adulterants in local supplies, which the agency says can help save lives.
- research
Apple Introduces Normalizing Trajectory Models for Few-Step Generation
Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as an expressive conditional normalizing flow with exact likelihood training. NTM combines shallow invertible blocks per step with a deep parallel architecture, addressing few-step generation without sacrificing the likelihood framework that distillation, consistency training, or adversarial methods discard.
- research
NVIDIA teaches robots to assemble GB300 tester trays
NVIDIA describes how it taught robots to assemble tester trays for its Grace Blackwell GB300 superchip, which powers AI training and inference. The work required skilled physical labor in factories worldwide. NVIDIA says the project offered lessons in robot learning, mechanical intelligence, and traditional engineering.
- research
Ai2's Bolmo byte-level language model technique published in Nature
Ai2 announced that the technique behind Bolmo, its fully open byte-level language models, has been published in Nature. New checkpoints show the approach generalizes beyond the Olmo model family to other model families, according to the Allen Institute for AI.