arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs—generated code, hypotheses,…
Tag: cs.AI updates on arXiv.org
Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks
arXiv:2609.09233v1 Announce Type: new Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon…
Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
arXiv:2609.09306v1 Announce Type: new Abstract: This paper investigates the hypothesis that the first-order structure of physical interactions, i.e.…
Adaptive Entangled Game Modules in Artificial General Intelligence
arXiv:2609.09226v1 Announce Type: new Abstract: We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive…
Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models
arXiv:2609.05779v1 Announce Type: cross Abstract: Large language models used for code editing can be trained and deployed in at least two output regimes:…
Closed-Loop Evaluation of Bird’s-Eye-View Maps from Cross-View Transformers as Inputs to Behavior-Cloning Policies
arXiv:2609.05783v1 Announce Type: cross Abstract: In autonomous driving, Bird’s-Eye View (BEV) representations provide a structured, top-down abstraction…
Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding
arXiv:2609.05764v1 Announce Type: cross Abstract: The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM)…
Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora
arXiv:2609.05766v1 Announce Type: cross Abstract: The dominant approach to building domain-specific pretraining corpora is to filter large web archives…
Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks
arXiv:2609.05794v1 Announce Type: cross Abstract: Open-weight large language models face a low-cost white-box threat from representation engineering…
SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation
arXiv:2609.05742v1 Announce Type: cross Abstract: American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data…
