arXiv:2609.24663v1 Announce Type: new Abstract: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills,…
Category: cs.AI updates on arXiv.org
Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents
arXiv:2609.24755v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This…
World State Generator
arXiv:2609.24744v1 Announce Type: new Abstract: Language agents solve complex tasks through plans and actions. A single step the world refuses puts the…
Construting Reverse Thinking: Developing Large Language Models’ Reverse Thingking Ability
arXiv:2609.24760v1 Announce Type: new Abstract: When facing complex problems, humans tend to try various ideas for different issues. Human thinking…
TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction
arXiv:2609.24677v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series…
Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis
arXiv:2609.24620v1 Announce Type: new Abstract: Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware…
The Endless Exam: Mathematical Constructions from Today’s Models toward Superintelligence
arXiv:2609.24555v1 Announce Type: new Abstract: We introduce the Endless Exam, a benchmark for measuring mathematical progress from today’s models toward…
Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging
arXiv:2609.24517v1 Announce Type: new Abstract: Model merging aims to combine multiple fine-tuned models derived from a common pretrained model into a…
DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
arXiv:2609.24662v1 Announce Type: new Abstract: LLM-based agents increasingly operate in environments where they interact with users, tools, and external…
Custom Named Entity Recognition and Topic Classification for Global Health Publications
arXiv:2609.24625v1 Announce Type: new Abstract: How should natural language processing models be selected and adapted for global health literature in…
