Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM,…
Author: script
AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories
arXiv:2609.05837v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training…
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
arXiv:2609.05818v1 Announce Type: new Abstract: We introduce ABLE, a benchmark for evaluating LLM agents’ ability to use biological AI models (BAIMs),…
Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment
arXiv:2609.05800v1 Announce Type: new Abstract: Activation steering controls LLM behavior at inference time by adding learned directions to hidden states,…
Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration
arXiv:2609.05801v1 Announce Type: new Abstract: A document modeled as a discrete sequence of tokens can be thought of as being generated from a…
More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review
arXiv:2609.05788v1 Announce Type: new Abstract: Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an…
Exposing Weaknesses in Emotion Recognition in Conversations
arXiv:2609.05806v1 Announce Type: new Abstract: Emotion Recognition in Conversations (ERC) aims to identify speakers’ emotions in multi-turn dialogue.…
AI News Brief Hourly Summary 2026-09-10 09h : 10 posts
10 posts published in the last hour 06:32Inference-Time Graph Engineering for Multi-Agent LLM Workflows 06:32DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents 06:32From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale 06:32The Normalization of…
Inference-Time Graph Engineering for Multi-Agent LLM Workflows
arXiv:2609.05774v1 Announce Type: new Abstract: Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate…
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
arXiv:2609.05776v1 Announce Type: new Abstract: Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely…
