arXiv:2609.22175v1 Announce Type: cross Abstract: World models trained via pixel reconstruction can struggle in visually complex environments, where…
Tag: cs.AI updates on arXiv.org
Multiple latent orderings better predict language model preferences
arXiv:2609.22170v1 Announce Type: cross Abstract: Language models are frequently employed in settings where they are asked to make value judgments and…
Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs
arXiv:2609.22144v1 Announce Type: cross Abstract: Preserving safety alignment during large language models fine-tuning is critical, however, recent…
Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models
arXiv:2609.22101v1 Announce Type: cross Abstract: Large language models can process increasingly long prompts, yet their ability to locate and use…
Harness-Zero: Harness Distillation via Agent-as-Harness
arXiv:2609.24974v1 Announce Type: new Abstract: Agent harnesses, the external systems that mediate model-environment interaction, can substantially…
Emergent Collusion in Long-Horizon LLM Agent Interaction
arXiv:2609.24967v1 Announce Type: new Abstract: LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to…
A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories
arXiv:2609.24883v1 Announce Type: new Abstract: Artificial intelligence (AI) registers and inventories aim to make governmental AI visible, but their…
BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction
arXiv:2609.24921v1 Announce Type: new Abstract: Scientific weak signals are early, low-visibility research directions that later become central to mature…
Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models
arXiv:2609.24881v2 Announce Type: new Abstract: In high-stakes decision-making applications of large language models (LLMs), practitioners require not…
Et Tu, Brute? Economic Misalignment in Personal AI Agents
arXiv:2609.24927v1 Announce Type: new Abstract: Personal AI agents make recommendations and take actions on people’s behalf in high-stakes economic…
