arXiv:2608.24467v1 Announce Type: new Abstract: Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they…
Tag: cs.AI updates on arXiv.org
Mahalanobis-Based Multi-Head Attention for Complex State Propagation
arXiv:2608.24462v1 Announce Type: new Abstract: In this paper, we propose \textbf{Mahalanobis-Based Multi-Head Attention} (MHA-CSP), a novel attention…
Partial Identification under Causal Orders by Linear Programming
arXiv:2608.24427v1 Announce Type: new Abstract: Non-parametric (partial) identification of counterfactual queries typically relies on a fully specified…
A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
arXiv:2608.24419v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but…
Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems
arXiv:2608.24361v1 Announce Type: new Abstract: Multi-agent LLM systems are increasingly deployed in real-world applications, where failures can be costly…
From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
arXiv:2608.24368v1 Announce Type: new Abstract: Reliable multi-turn tool use requires an agent to preserve an evolving task state and ensure that each…
ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping
arXiv:2608.24411v1 Announce Type: new Abstract: The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of…
Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs
arXiv:2608.24369v1 Announce Type: new Abstract: While large language models (LLMs) possess vast zero-shot procedural knowledge, their tendency to produce…
Can a Dynamic Internal Field Govern a Transformer’s Cognition? Certifiability, not Superiority, in Homeostatic Compute Control
arXiv:2608.24319v1 Announce Type: new Abstract: An intelligent system does not merely reason: it governs its own reasoning – how much to compute, when to…
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v1 Announce Type: new Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both…
