PyTorch Optimizes Jagged Flash Attention with TLX Toward SOTA FA4 on Blackwell
PyTorch published work on Jagged Flash Attention (JFA), the attention kernel behind Meta's Generative Ads Model (GEM), running on NVIDIA Blackwell (B200). The post details using TLX to optimize the kernel, describing the effort as a road toward state-of-the-art Flash Attention 4 performance on Blackwell hardware.
Sources and evidence
Summary last validated Oct 2, 2026
Reader actions
Report an issue
Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.
Related coverage
- infrastructure
PyTorch details session-aware agentic inference with NVIDIA Dynamo
PyTorch published a blog post describing session-aware agentic inference with NVIDIA Dynamo. It explains that agentic workloads differ from single-turn chat: an agent session can include a large initial prefill, repeated model calls, and parallel subagents, changing the traffic an inference server handles.
- open source
PyTorch adds Spyre as native device via torch-spyre
PyTorch's blog details how Spyre becomes a native PyTorch device. The integration connects PyTorch's existing device, allocator, stream, and compiler abstractions through torch-spyre to the Spyre runtime and firmware. PrivateUse1 is used to give Spyre a real device registration within PyTorch.
- research
NIST Study Maps How Toxic Adulterants in Fentanyl Vary Across US
NIST researchers analyzed fentanyl samples nationwide and found that the toxic substances mixed into the drug vary by region and shift over time. The study aims to help first responders and law enforcement identify dangerous adulterants in local supplies, which the agency says can help save lives.
- research
Apple Introduces Normalizing Trajectory Models for Few-Step Generation
Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as an expressive conditional normalizing flow with exact likelihood training. NTM combines shallow invertible blocks per step with a deep parallel architecture, addressing few-step generation without sacrificing the likelihood framework that distillation, consistency training, or adversarial methods discard.
- research
NVIDIA teaches robots to assemble GB300 tester trays
NVIDIA describes how it taught robots to assemble tester trays for its Grace Blackwell GB300 superchip, which powers AI training and inference. The work required skilled physical labor in factories worldwide. NVIDIA says the project offered lessons in robot learning, mechanical intelligence, and traditional engineering.