arXiv:2609.09657v1 Announce Type: new Abstract: Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions…
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
arXiv:2609.09647v1 Announce Type: new Abstract: Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real…
ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance
arXiv:2609.09458v1 Announce Type: new Abstract: As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather…
CityPlanner: A Sandbox Agent for Executable Urban Planning
arXiv:2609.09578v1 Announce Type: new Abstract: Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from…
From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins
arXiv:2609.09625v1 Announce Type: new Abstract: As Digital Twin (DT) systems evolve beyond state synchronization toward task-oriented and knowledge-driven…
Multi-Agent Agentic Graph Learning via Structural Signatures
arXiv:2609.09565v1 Announce Type: new Abstract: Agentic graph learning (AGL) has recently achieved promising results on graph reasoning tasks, where an…
A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
arXiv:2609.09589v1 Announce Type: new Abstract: Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast…
AI News Brief Hourly Summary 2026-09-11 07h : 10 posts
10 posts published in the last hour 04:32Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations 04:32Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration 04:32XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?…
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
arXiv:2609.09448v1 Announce Type: new Abstract: As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the…
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
arXiv:2609.09418v1 Announce Type: new Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated…
XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
arXiv:2609.09428v1 Announce Type: new Abstract: Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging…
The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
arXiv:2609.09395v1 Announce Type: new Abstract: Language models act through tools, yet practical agents face libraries containing thousands of interfaces.…
Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery
arXiv:2609.09413v1 Announce Type: new Abstract: Choosing a recovery process for scale-up requires connecting laboratory results with product requirements,…
An Autonomous GeoAI Agent for Arctic Eco-Navigation
arXiv:2609.09374v1 Announce Type: new Abstract: Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand…
OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs—generated code, hypotheses,…
Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks
arXiv:2609.09233v1 Announce Type: new Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon…
Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
arXiv:2609.09306v1 Announce Type: new Abstract: This paper investigates the hypothesis that the first-order structure of physical interactions, i.e.…
Adaptive Entangled Game Modules in Artificial General Intelligence
arXiv:2609.09226v1 Announce Type: new Abstract: We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive…
