AI News Brief: today roundup Researchers discovered that replacing standard 0–100 confidence scales with a 0–20 format significantly improves how accurately LLMs express uncertainty. Researchers introduced VeriSim, an open-source evaluation framework that stress-tests medical LLMs against realistic patient communication noise.…
Read more →