arXiv:2605.27752v3 Announce Type: replace Abstract: Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token…
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
arXiv:2606.08531v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into…
Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation
arXiv:2606.06076v3 Announce Type: replace Abstract: While Vision-Language Models excel at general multimodal understanding, they still struggle with…
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
arXiv:2607.19949v4 Announce Type: replace Abstract: Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires…
DATAREEL: Automated Data-Driven Video Story Generation with Animations
arXiv:2604.25220v2 Announce Type: replace Abstract: Data videos combine animated visualizations with synchronized narration to communicate quantitative…
In-Context Examples Suppress Scientific Knowledge Recall in LLMs
arXiv:2604.27540v2 Announce Type: replace Abstract: Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden…
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On
arXiv:2605.19035v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable…
Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?
arXiv:2605.22148v3 Announce Type: replace Abstract: A large language model (LLM) agent that writes and edits its own skill library must also decide which…
Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
arXiv:2605.00123v3 Announce Type: replace Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through…
AI News Brief Hourly Summary 2026-08-11 03h : 13 posts
13 posts were published in the last hour 0:32 : Counterfactual Simulation Training for Chain-of-Thought Faithfulness 0:32 : AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation 0:32 : MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large…
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
arXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM…
AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation
arXiv:2603.20986v2 Announce Type: replace Abstract: Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics…
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
arXiv:2604.16009v2 Announce Type: replace Abstract: Most large language model benchmarks evaluate final-answer quality but reveal little about how models…
Alignment has a Fantasia Problem
arXiv:2604.21827v2 Announce Type: replace Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g.,…
OpenAI reportedly completed a $7 billion employee tender offer
San Francisco’s housing market is in trouble again.
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
arXiv:2603.21607v2 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it…
Social World Models
arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about…
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
arXiv:2608.07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their…
