arXiv:2609.22609v1 Announce Type: cross Abstract: Insertion is a fundamental operation in robotic construction assembly, where variations in material…
Category: AI
Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that…
Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
arXiv:2609.22588v1 Announce Type: cross Abstract: Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make…
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
arXiv:2609.22586v1 Announce Type: cross Abstract: Modern audio-language models are no longer judged only on what words they can transcribe, but on whether…
Do Student LLMs Inherit OOD Robustness? Invariance-Weighted Distillation for Reliable Knowledge Transfer
arXiv:2609.22566v1 Announce Type: cross Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students.…
FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation
arXiv:2609.22538v1 Announce Type: cross Abstract: Large language model (LLM) planners can decompose natural-language instructions and select reusable…
The Ups and Downs of Backprop Weights
arXiv:2609.22554v1 Announce Type: cross Abstract: Backpropagation (BP) has driven the remarkable success of modern deep learning by enabling large…
Beyond the Leaderboard: Counterfactual Diagnosis of End-to-End and VLA Driving Policies Under Domain Shift
arXiv:2609.22582v1 Announce Type: cross Abstract: End-to-end and vision-language-action (VLA) driving policies are compared by leaderboard rank, but a…
Sam Altman’s remarks at the United Nations Security Council
OpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.
Zero-Trust Authorization and Discovery for Enterprise MCP
arXiv:2609.22573v1 Announce Type: cross Abstract: LLM agents translate natural-language context, which may include attacker-controlled text, into…
