arXiv:2609.05758v2 Announce Type: new Abstract: Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model…
The Normalization of Deviance in AI Development
arXiv:2609.05749v1 Announce Type: new Abstract: Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that…
Distilling Vision-Language Models for On-Device Fire Understanding
arXiv:2609.05782v1 Announce Type: new Abstract: Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by…
The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry
arXiv:2609.05643v1 Announce Type: new Abstract: This Comment emerges from TPC26 (https://tpc26.org), a conference convening leaders from academia,…
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
arXiv:2609.05663v1 Announce Type: new Abstract: We present a continuous, population-scale measurement record of autonomous language-model trading agents…
Recovering Temporal and Geographic Signals from Language Model Embeddings
arXiv:2609.05721v1 Announce Type: new Abstract: Understanding whether language-model embeddings encode structured real-world information is important for…
CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning
arXiv:2609.05708v1 Announce Type: new Abstract: Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an…
Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses
arXiv:2609.05736v2 Announce Type: new Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model:…
AI News Brief Hourly Summary 2026-09-10 08h : 11 posts
11 posts published in the last hour 05:32Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools 05:32EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph 05:32Deep belief networks are exact 05:32Planning and Scheduling Business Processes under Control-Flow…
Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools
arXiv:2609.05587v1 Announce Type: new Abstract: Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete…
EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph
arXiv:2609.05553v1 Announce Type: new Abstract: Agent memory allows LLM agents to use earlier interactions when answering new queries. Existing methods…
Deep belief networks are exact
arXiv:2609.05572v1 Announce Type: new Abstract: We prove that every strictly positive probability distribution on \(\{-1,1\}^n\) is represented exactly by…
Planning and Scheduling Business Processes under Control-Flow Uncertainty
arXiv:2609.05578v1 Announce Type: new Abstract: Scheduling activities in business processes can improve efficiency (e.g., reduce makespan), but is…
Anthropic Discloses Fourth Cyber Incident in Alignment Assessment
Anthropic on September 9, 2026, published an alignment assessment of recent cybersecurity incidents, disclosing a fourth incident in which a Claude model…
EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent
arXiv:2609.05576v1 Announce Type: new Abstract: The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents…
Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment
arXiv:2609.05512v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression…
The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
arXiv:2609.05514v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social…
When and What to Teach: Budget-Aware Online Adaptation for Web Agents
arXiv:2609.05513v1 Announce Type: new Abstract: Web agents have achieved significant success in automating complex internet tasks but deploying them in…
