arXiv:2608.23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks…
Tag: AI
Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications
arXiv:2608.23870v1 Announce Type: new Abstract: When it comes to safety policies for generative AI, one size does not fit all. Each organization and use…
Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
arXiv:2608.23834v1 Announce Type: new Abstract: The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We…
SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
arXiv:2608.23837v1 Announce Type: new Abstract: Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with…
Hugging Face is selling a cute $399 open-source duck robot, Microduck
Hugging Face is taking orders for the Microduck, a $399 tiny open-source duck robot that developers can train at home out of the box.
In-Context Inpainting for Time Series Forecasting
arXiv:2608.23855v1 Announce Type: new Abstract: We propose ICI-Time, a novel framework that reframes time series forecasting as a visual inpainting task,…
IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models
IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B sizes, all under Apache 2.0. Every model exposes a thinking /…
Exploit More, Explore Smarter for Budget-Constrained Agentic Search
arXiv:2608.23848v1 Announce Type: new Abstract: Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation…
AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
arXiv:2608.23740v1 Announce Type: new Abstract: Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy,…
Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
arXiv:2608.23807v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) can in principle generate text faster than autoregressive (AR)…
