arXiv:2608.24419v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but…
Category: AI
Agent Washing: Why Some Restaurant Operators Are Wary of Overhyped AI
For those in the restaurant industry, it seems like AI rollouts have hit a rocky road. In the Spring, Starbucks switched off a computer-vision system for…
ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping
arXiv:2608.24411v1 Announce Type: new Abstract: The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of…
Vanguard to Acquire AI Custody Platform Altruist
Vanguard has agreed to acquire Altruist, an AI-forward wealth technology and custody platform for independent financial advisors, the two companies…
Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems
arXiv:2608.24361v1 Announce Type: new Abstract: Multi-agent LLM systems are increasingly deployed in real-world applications, where failures can be costly…
Can a Dynamic Internal Field Govern a Transformer’s Cognition? Certifiability, not Superiority, in Homeostatic Compute Control
arXiv:2608.24319v1 Announce Type: new Abstract: An intelligent system does not merely reason: it governs its own reasoning – how much to compute, when to…
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v1 Announce Type: new Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both…
SonarLLM: A Native Sonar–Optical Multimodal Large Language Model for Underwater Perception
arXiv:2608.24325v1 Announce Type: new Abstract: Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras…
Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown
Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables,…
Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning
arXiv:2608.24338v1 Announce Type: new Abstract: Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, yet…
