arXiv:2505.22451v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have made significant progress in mathematical capabilities in recent…
Tag: AI
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
arXiv:2609.02771v1 Announce Type: cross Abstract: Training data attribution (TDA) aims to identify training examples that shape model behavior, but its…
HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design
arXiv:2609.02746v1 Announce Type: cross Abstract: Polymeric materials are central to modern technologies, with applications ranging from energy to health…
Untangling the Mechanisms of Misleading Context in Medical Question Answering
arXiv:2609.02754v1 Announce Type: cross Abstract: Large language models now answer medical questions with expert-level performance. However, the context…
GPT-6 Astra is the first model making OpenAI willing to declare the “AGI era”
OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the “AGI era.” Astra tops benchmarks in…
frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study
arXiv:2609.02804v1 Announce Type: cross Abstract: For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its…
AI Efficiency Could Cost Us the Next Generation of Experts
A little over a decade ago, I led the controls design for a first-of-its-kind full digital-control system for a U.S. nuclear plant. It was, on paper, a…
Dutch Books for Language Models
arXiv:2609.02797v1 Announce Type: cross Abstract: People increasingly use language models to support life decisions. Many such decisions involve a…
Language Models Can Control Their Own Attention
arXiv:2609.02737v1 Announce Type: cross Abstract: Language models spend most of their attention on a small fraction of context, yet they read the entire…
RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models
arXiv:2609.02731v1 Announce Type: cross Abstract: Large vision-language models have achieved remarkable success in vision-language tasks. However, they…
