arXiv:2608.23873v1 Announce Type: new Abstract: Everything a language model sees is tokens. The serving stack knows what each span is — user input, tool…
Tag: cs.AI updates on arXiv.org
Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems
arXiv:2608.23906v1 Announce Type: new Abstract: Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including…
BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification
arXiv:2608.23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks…
Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications
arXiv:2608.23870v1 Announce Type: new Abstract: When it comes to safety policies for generative AI, one size does not fit all. Each organization and use…
Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
arXiv:2608.23834v1 Announce Type: new Abstract: The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We…
SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
arXiv:2608.23837v1 Announce Type: new Abstract: Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with…
In-Context Inpainting for Time Series Forecasting
arXiv:2608.23855v1 Announce Type: new Abstract: We propose ICI-Time, a novel framework that reframes time series forecasting as a visual inpainting task,…
Exploit More, Explore Smarter for Budget-Constrained Agentic Search
arXiv:2608.23848v1 Announce Type: new Abstract: Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation…
AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
arXiv:2608.23740v1 Announce Type: new Abstract: Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy,…
Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
arXiv:2608.23807v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) can in principle generate text faster than autoregressive (AR)…
