arXiv:2609.19636v1 Announce Type: new Abstract: Reinforcement learning now trains language-model agents that act over dozens of steps in live…
Tag: cs.AI updates on arXiv.org
SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership
arXiv:2609.19610v1 Announce Type: new Abstract: Understanding humans over long horizons requires agents to infer not only what people need in the moment,…
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
arXiv:2609.19644v1 Announce Type: new Abstract: Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and…
From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization
arXiv:2609.19630v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking…
Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks
arXiv:2609.19538v1 Announce Type: new Abstract: Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned…
EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
arXiv:2609.19523v1 Announce Type: new Abstract: Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier…
Self Improvement via Fast Tree-search
arXiv:2609.19526v1 Announce Type: new Abstract: Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While…
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
arXiv:2609.19524v1 Announce Type: new Abstract: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial…
When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\’esum\’e Screening
arXiv:2609.19530v1 Announce Type: new Abstract: Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their…
An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
arXiv:2609.19519v1 Announce Type: new Abstract: Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an…
