arXiv:2609.05992v1 Announce Type: new Abstract: As LLM-based agents continue to advance, their evaluation has become increasingly multifaceted: a capable…
Tag: AI
Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review
arXiv:2609.05947v1 Announce Type: new Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer…
Learning Counterfactual World Models for Embodied Reasoning under Partial Observability
arXiv:2609.05834v1 Announce Type: new Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then…
Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization
arXiv:2609.05889v1 Announce Type: new Abstract: Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal…
The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble
arXiv:2609.05894v1 Announce Type: new Abstract: The exponentiation of Artificial intelligence (AI) in the recent past has entered a transformative era…
Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
arXiv:2609.05824v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external skills, but routing user requests over…
DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM,…
AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories
arXiv:2609.05837v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training…
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
arXiv:2609.05818v1 Announce Type: new Abstract: We introduce ABLE, a benchmark for evaluating LLM agents’ ability to use biological AI models (BAIMs),…
Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment
arXiv:2609.05800v1 Announce Type: new Abstract: Activation steering controls LLM behavior at inference time by adding learned directions to hidden states,…
