arXiv:2609.05801v1 Announce Type: new Abstract: A document modeled as a discrete sequence of tokens can be thought of as being generated from a…
Category: AI
More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review
arXiv:2609.05788v1 Announce Type: new Abstract: Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an…
Exposing Weaknesses in Emotion Recognition in Conversations
arXiv:2609.05806v1 Announce Type: new Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers’ emotions in multi-turn dialogue.…
Inference-Time Graph Engineering for Multi-Agent LLM Workflows
arXiv:2609.05774v1 Announce Type: new Abstract: Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate…
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
arXiv:2609.05776v1 Announce Type: new Abstract: Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely…
From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale
arXiv:2609.05758v2 Announce Type: new Abstract: Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model…
The Normalization of Deviance in AI Development
arXiv:2609.05749v1 Announce Type: new Abstract: Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that…
Distilling Vision-Language Models for On-Device Fire Understanding
arXiv:2609.05782v1 Announce Type: new Abstract: Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by…
The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry
arXiv:2609.05643v1 Announce Type: new Abstract: This Comment emerges from TPC26 (https://tpc26.org), a conference convening leaders from academia,…
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
arXiv:2609.05663v1 Announce Type: new Abstract: We present a continuous, population-scale measurement record of autonomous language-model trading agents…
