arXiv:2609.05758v2 Announce Type: new Abstract: Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model…
Author: script
The Normalization of Deviance in AI Development
arXiv:2609.05749v1 Announce Type: new Abstract: Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that…
Distilling Vision-Language Models for On-Device Fire Understanding
arXiv:2609.05782v1 Announce Type: new Abstract: Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by…
The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry
arXiv:2609.05643v1 Announce Type: new Abstract: This Comment emerges from TPC26 (https://tpc26.org), a conference convening leaders from academia,…
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
arXiv:2609.05663v1 Announce Type: new Abstract: We present a continuous, population-scale measurement record of autonomous language-model trading agents…
Recovering Temporal and Geographic Signals from Language Model Embeddings
arXiv:2609.05721v1 Announce Type: new Abstract: Understanding whether language-model embeddings encode structured real-world information is important for…
CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning
arXiv:2609.05708v1 Announce Type: new Abstract: Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an…
Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses
arXiv:2609.05736v2 Announce Type: new Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model:…
AI News Brief Hourly Summary 2026-09-10 08h : 11 posts
11 posts published in the last hour 05:32Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools 05:32EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph 05:32Deep belief networks are exact 05:32Planning and Scheduling Business Processes under Control-Flow…
Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools
arXiv:2609.05587v1 Announce Type: new Abstract: Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete…
