arXiv:2609.06036v1 Announce Type: new Abstract: Proposal-based controllers—learned policies, language-model planners, and other black-box…
Tag: cs.AI updates on arXiv.org
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
arXiv:2609.06059v1 Announce Type: new Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation…
MOAE: Multi-Objective Agent Evolution with Pareto-Preserving Search
arXiv:2609.05992v1 Announce Type: new Abstract: As LLM-based agents continue to advance, their evaluation has become increasingly multifaceted: a capable…
Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review
arXiv:2609.05947v1 Announce Type: new Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer…
Learning Counterfactual World Models for Embodied Reasoning under Partial Observability
arXiv:2609.05834v1 Announce Type: new Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then…
Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization
arXiv:2609.05889v1 Announce Type: new Abstract: Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal…
The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble
arXiv:2609.05894v1 Announce Type: new Abstract: The exponentiation of Artificial intelligence (AI) in the recent past has entered a transformative era…
Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
arXiv:2609.05824v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external skills, but routing user requests over…
AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories
arXiv:2609.05837v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training…
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
arXiv:2609.05818v1 Announce Type: new Abstract: We introduce ABLE, a benchmark for evaluating LLM agents’ ability to use biological AI models (BAIMs),…
