arXiv:2604.21827v2 Announce Type: replace Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g.,…
Category: cs.AI updates on arXiv.org
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
arXiv:2603.21607v2 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it…
Social World Models
arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about…
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
arXiv:2608.07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their…
“LLM Agent Performance” Is Not a Single Evaluation Target
arXiv:2602.03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness,…
Boundary Density Likelihood for Direct Event-Time Supervision
arXiv:2408.12792v2 Announce Type: replace Abstract: Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models…
Serious Games: Human-AI Interaction, Evolution, and Coevolution
arXiv:2505.16388v3 Announce Type: replace Abstract: The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models…
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making…
Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools
arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational,…
CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
arXiv:2608.07458v1 Announce Type: cross Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache…
