arXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM…
Category: cs.AI updates on arXiv.org
AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation
arXiv:2603.20986v2 Announce Type: replace Abstract: Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics…
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
arXiv:2604.16009v2 Announce Type: replace Abstract: Most large language model benchmarks evaluate final-answer quality but reveal little about how models…
Alignment has a Fantasia Problem
arXiv:2604.21827v2 Announce Type: replace Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g.,…
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
arXiv:2603.21607v2 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it…
Social World Models
arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about…
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
arXiv:2608.07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their…
“LLM Agent Performance” Is Not a Single Evaluation Target
arXiv:2602.03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness,…
Boundary Density Likelihood for Direct Event-Time Supervision
arXiv:2408.12792v2 Announce Type: replace Abstract: Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models…
Serious Games: Human-AI Interaction, Evolution, and Coevolution
arXiv:2505.16388v3 Announce Type: replace Abstract: The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models…