arXiv:2609.22601v1 Announce Type: cross Abstract: The increasing reliance of autonomous AI agents on external and distributed knowledge sources introduces…
Tag: cs.AI updates on arXiv.org
From Documented Strengths to Force Limits: Material-Informed Robotic Insertion for Construction Assembly
arXiv:2609.22609v1 Announce Type: cross Abstract: Insertion is a fundamental operation in robotic construction assembly, where variations in material…
Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that…
Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
arXiv:2609.22588v1 Announce Type: cross Abstract: Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make…
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
arXiv:2609.22586v1 Announce Type: cross Abstract: Modern audio-language models are no longer judged only on what words they can transcribe, but on whether…
Do Student LLMs Inherit OOD Robustness? Invariance-Weighted Distillation for Reliable Knowledge Transfer
arXiv:2609.22566v1 Announce Type: cross Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students.…
FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation
arXiv:2609.22538v1 Announce Type: cross Abstract: Large language model (LLM) planners can decompose natural-language instructions and select reusable…
The Ups and Downs of Backprop Weights
arXiv:2609.22554v1 Announce Type: cross Abstract: Backpropagation (BP) has driven the remarkable success of modern deep learning by enabling large…
Beyond the Leaderboard: Counterfactual Diagnosis of End-to-End and VLA Driving Policies Under Domain Shift
arXiv:2609.22582v1 Announce Type: cross Abstract: End-to-end and vision-language-action (VLA) driving policies are compared by leaderboard rank, but a…
Zero-Trust Authorization and Discovery for Enterprise MCP
arXiv:2609.22573v1 Announce Type: cross Abstract: LLM agents translate natural-language context, which may include attacker-controlled text, into…
