arXiv:2609.28984v1 Announce Type: cross Abstract: Cross-embodiment learning has shown that a single model, such as a vision-language-action (VLA) model,…
Tag: cs.AI updates on arXiv.org
Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
arXiv:2609.28991v1 Announce Type: cross Abstract: Video understanding is increasingly performed by multi-stage LLM agents that separate temporal…
Cross-Country Code-Mixing for Generative Recommendation
arXiv:2609.28972v1 Announce Type: cross Abstract: Cross-country recommendation on modern e-commerce platforms is typically deployed with disjoint user and…
Calibrated Decision Models for Autonomous Penetration-Testing Harnesses: JEV and Laya as System One Decision Layers for LLM-Driven Pentest Agents
arXiv:2609.28940v1 Announce Type: cross Abstract: Autonomous penetration-testing harnesses use large language models (LLMs) for reconnaissance,…
Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots
arXiv:2609.29043v1 Announce Type: cross Abstract: General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to…
Blockchain-Enabled Artificial Intelligence and AI Agents for Secure Data Sharing and Cybersecurity Applications
arXiv:2609.28843v1 Announce Type: cross Abstract: Blockchain and artificial intelligence (AI) are converging into a single infrastructural layer for…
Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding
arXiv:2609.28854v1 Announce Type: cross Abstract: Language-model agents increasingly answer questions over customer-relationship management (CRM) records,…
On the Effectiveness of Kernel-Level Evidence for Agent Security
arXiv:2609.28915v1 Announce Type: cross Abstract: LLM agents are deployed into infrastructure that grants them broad host authority, yet existing…
Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots
arXiv:2609.28910v1 Announce Type: cross Abstract: Effective robot assistance beyond narrow roles and repetitive tasks requires robots to be proactive – to…
Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs
arXiv:2609.28879v1 Announce Type: cross Abstract: Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty…
