arXiv:2609.02168v1 Announce Type: new Abstract: Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular…
Tag: cs.AI updates on arXiv.org
Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning
arXiv:2609.02191v1 Announce Type: new Abstract: Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We…
Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics
arXiv:2609.02116v1 Announce Type: new Abstract: Reverse-logistics operators often decide how to inspect and route returned assets before their condition…
EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision
arXiv:2609.02133v1 Announce Type: new Abstract: Empathetic response generation requires models to decide not only what to say, but also how to respond to…
MASkills: Continual Skills Optimization for Multi-Agent LLM Systems
arXiv:2609.02094v1 Announce Type: new Abstract: LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement…
CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning
arXiv:2609.02074v1 Announce Type: new Abstract: Planning is a central capability that enables agents to decompose complex long-horizon tasks into…
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
arXiv:2609.02067v1 Announce Type: new Abstract: Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another’s…
READY or Not: Reliable Enterprise Agent Deployment
arXiv:2609.02095v1 Announce Type: new Abstract: An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent…
Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems
arXiv:2609.02092v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet…
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual…
