arXiv:2608.24419v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but…
Tag: cs.AI updates on arXiv.org
ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping
arXiv:2608.24411v1 Announce Type: new Abstract: The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of…
Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems
arXiv:2608.24361v1 Announce Type: new Abstract: Multi-agent LLM systems are increasingly deployed in real-world applications, where failures can be costly…
Can a Dynamic Internal Field Govern a Transformer’s Cognition? Certifiability, not Superiority, in Homeostatic Compute Control
arXiv:2608.24319v1 Announce Type: new Abstract: An intelligent system does not merely reason: it governs its own reasoning – how much to compute, when to…
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v1 Announce Type: new Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both…
SonarLLM: A Native Sonar–Optical Multimodal Large Language Model for Underwater Perception
arXiv:2608.24325v1 Announce Type: new Abstract: Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras…
Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning
arXiv:2608.24338v1 Announce Type: new Abstract: Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, yet…
The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
arXiv:2608.24358v1 Announce Type: new Abstract: Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As…
OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning
arXiv:2608.24310v1 Announce Type: new Abstract: Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from…
Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling
arXiv:2608.24274v1 Announce Type: new Abstract: A sustainable diet represents a multi-dimensional synergy among four essential pillars: nutrition…
