arXiv:2609.19124v1 Announce Type: new Abstract: Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key…
Category: AI
Why Infrastructure AI — Not Consumer AI — Will Define the Next Decade
In the last few years, the conversation around AI has accelerated. Each week brings a new AI tool to the scene, along with promises about how it will…
MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
arXiv:2609.19088v1 Announce Type: new Abstract: Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their…
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
arXiv:2609.18842v1 Announce Type: new Abstract: The scaling laws hold that a language model grows more capable with more parameters and more training…
Lost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning
arXiv:2609.18991v1 Announce Type: new Abstract: Multimodal LLMs report strong performance on scientific reasoning benchmarks, yet most treat perception…
Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows
arXiv:2609.18820v1 Announce Type: new Abstract: Agentic workflows now make consequential decisions in regulated settings, and the governance placed around…
Tower and NewPhotonics Ship Laser-Integrated PICs for AI Interconnect
Tower Semiconductor and NewPhotonics have begun high-volume shipments of laser-integrated, serviceable optical engine photonic integrated circuits for…
Suppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing
arXiv:2609.18985v1 Announce Type: new Abstract: Knowledge-editing benchmarks certify local correctness, whether an edited model produces the new fact on…
Adecco Group rolls out Agentforce Coworker to 27,000 staff in 40-plus countries
The Adecco Group is rolling out Salesforce’s Agentforce Coworker across more than 40 countries following a pilot in the UK and France, the staffing group…
Function Lives Where Variance Doesn’t: Task-Weighted Charts of a Language Model’s Computation
arXiv:2609.18989v1 Announce Type: new Abstract: How many dimensions does a language model’s computation actually use? The question is ill-posed until one…
