CapQuiz benchmark evaluates video captions via multiple-choice questions
Researchers introduce CapQuiz, a reference-free benchmark that scores video captions by how well they answer human-verified multiple-choice questions derived from each video. It uses a hierarchical taxonomy of 10 question types across 24 video domains. Experiments show CapQuiz correlates better with human judgments than existing metrics and gives interpretable, fine-grained insights into Visual Large Language Model caption quality.
Why this appeared
- Multiple sources
These labels describe how the story was ranked for the feed — not editorial endorsements.
Sources and evidence
Summary last validated Sep 11, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- research
Ai2 presents Olmo Hybrid, Asta, and Bolmo at COLM 2026
Ai2 attended COLM 2026 to present its work on open language models, training infrastructure, and AI for science. The organization highlighted Olmo Hybrid, Asta, and Bolmo, and connected with the broader research community at the conference.
- research
NIST Study Maps How Toxic Adulterants in Fentanyl Vary Across US
NIST researchers analyzed fentanyl samples nationwide and found that the toxic substances mixed into the drug vary by region and shift over time. The study aims to help first responders and law enforcement identify dangerous adulterants in local supplies, which the agency says can help save lives.
- research
Apple Introduces Normalizing Trajectory Models for Few-Step Generation
Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as an expressive conditional normalizing flow with exact likelihood training. NTM combines shallow invertible blocks per step with a deep parallel architecture, addressing few-step generation without sacrificing the likelihood framework that distillation, consistency training, or adversarial methods discard.
- research
NVIDIA teaches robots to assemble GB300 tester trays
NVIDIA describes how it taught robots to assemble tester trays for its Grace Blackwell GB300 superchip, which powers AI training and inference. The work required skilled physical labor in factories worldwide. NVIDIA says the project offered lessons in robot learning, mechanical intelligence, and traditional engineering.
- research
Ai2's Bolmo byte-level language model technique published in Nature
Ai2 announced that the technique behind Bolmo, its fully open byte-level language models, has been published in Nature. New checkpoints show the approach generalizes beyond the Olmo model family to other model families, according to the Allen Institute for AI.