research
34 sourcesResearch

New Research: LLM Judges Unstable, RAG Privacy, and More

New arXiv papers cover LLM judge instability (Wiggle Framework), privacy-preserving RAG (SEAG), multi-agent pricing, transformer expressiveness, constraint saturation, Dual-Flow Transformers, governed memory, and edge scheduling. Findings: judges flip verdicts 25-91% under pressure; SEAG hides sensitive entities with >74% accuracy; instruction following collapses beyond 5-6 constraints.

Why this appeared

  • Multiple sources

These labels describe how the story was ranked for the feed — not editorial endorsements.

Sources and evidence

Summary last validated Aug 15, 2026

Reader actions

Report an issue

Use this for an incorrect summary, wrong source, duplicate story, or wrong category. Submissions are private and do not change the story automatically.

New Research: LLM Judges Unstable, RAG Privacy, and More · NewsAI