arXiv:2609.14744v2 Announce Type: replace Abstract: By acquiring compute, credentials, accounts, services, and other agents, autonomous AI agents can…
Tag: cs.AI updates on arXiv.org
Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
arXiv:2609.18366v2 Announce Type: replace Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a…
Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
arXiv:2609.18461v2 Announce Type: replace Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit…
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
arXiv:2609.19524v2 Announce Type: replace Abstract: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern…
Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
arXiv:2608.26088v2 Announce Type: replace Abstract: Addressing critical global challenges, from food security and disaster risk to disease outbreaks and…
A visual large language foundational model for medical image recognition using clinician-contributed online resources
arXiv:2609.06914v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing…
Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
arXiv:2609.13519v2 Announce Type: replace Abstract: We introduce Fraglingo, an autoregressive molecular generator that constructs molecules step by step…
Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison
arXiv:2608.30044v3 Announce Type: replace Abstract: Model comparison increasingly relies on large collections of publicly reported benchmark scores, yet…
MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents
arXiv:2609.14399v2 Announce Type: replace Abstract: Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent…
Intent-Governed Tool Authorization for AI Agents
arXiv:2606.22916v4 Announce Type: replace Abstract: Tool-using AI agents commonly operate under integration credentials whose static permissions exceed a…
