arXiv:2609.24876v1 Announce Type: new Abstract: In multi-agent social settings, model reliability varies across relationships. Beyond inferring what…
Tag: cs.AI updates on arXiv.org
Convex AI Compositionality and the Governance of AI System Populations
arXiv:2609.24784v1 Announce Type: new Abstract: AI governance increasingly requires providers and public authorities to reason about multiple AI…
GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes
arXiv:2609.24831v1 Announce Type: new Abstract: Agents have attracted considerably increasing attention due to the power of executing both Reasoning and…
MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution
arXiv:2609.24838v1 Announce Type: new Abstract: Medical agents increasingly combine general reasoning models with specialized clinical tools, yet their…
Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection
arXiv:2609.24855v1 Announce Type: new Abstract: Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its…
Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents
arXiv:2609.24663v1 Announce Type: new Abstract: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills,…
Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents
arXiv:2609.24755v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This…
World State Generator
arXiv:2609.24744v1 Announce Type: new Abstract: Language agents solve complex tasks through plans and actions. A single step the world refuses puts the…
Construting Reverse Thinking: Developing Large Language Models’ Reverse Thingking Ability
arXiv:2609.24760v1 Announce Type: new Abstract: When facing complex problems, humans tend to try various ideas for different issues. Human thinking…
TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction
arXiv:2609.24677v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series…
