
Efficient Knowledge Distillation Cuts VRAM, Enables Single-GPU Training
Multiverse Computing and collaborators propose offline top-K logits caching and a fused chunked KL loss to make knowledge distillation of large LLMs cheaper. Their method avoids holding teacher and student simultaneously and never materializes full vocabulary-size matrices, reducing VRAM enough for single-GPU training and large-scale experimentation.
Sources and evidence
Summary last validated Aug 30, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- infrastructure
Ai2 replaces priority GPU scheduler with budget-based fair-share system
Ai2's AI Infrastructure team replaced its priority-based GPU scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and time-slicing contracts. The institute manages thousands of NVIDIA H100, B200, and B300 GPUs across 88- to 1024-GPU clusters for about 150 researchers, with demand running 2-3x available capacity. The change moved GPU allocation debates into a transparent administrative budgeting process.
- research
NIST Study Maps How Toxic Adulterants in Fentanyl Vary Across US
NIST researchers analyzed fentanyl samples nationwide and found that the toxic substances mixed into the drug vary by region and shift over time. The study aims to help first responders and law enforcement identify dangerous adulterants in local supplies, which the agency says can help save lives.
- research
Apple Introduces Normalizing Trajectory Models for Few-Step Generation
Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as an expressive conditional normalizing flow with exact likelihood training. NTM combines shallow invertible blocks per step with a deep parallel architecture, addressing few-step generation without sacrificing the likelihood framework that distillation, consistency training, or adversarial methods discard.
- research
NVIDIA teaches robots to assemble GB300 tester trays
NVIDIA describes how it taught robots to assemble tester trays for its Grace Blackwell GB300 superchip, which powers AI training and inference. The work required skilled physical labor in factories worldwide. NVIDIA says the project offered lessons in robot learning, mechanical intelligence, and traditional engineering.
- models
Liquid AI releases d1-3B and d1-omni-600M open decision models for edge
Liquid AI released two open decision models: d1-3B and experimental d1-omni-600M. d1-3B scores 48.57 on Decision Index 0.2.1, topping sub-10B models and Decider 35B-A3B, and answers in 16 ms on NVIDIA Jetson AGX Thor. d1-omni-600M handles text with image or audio. Both are built on Liquid Foundation Models.