arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return “no evidence” either because a benchmark is clean or because…
Category: cs.AI updates on arXiv.org
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing…
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic…
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be…
GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
arXiv:2608.07881v1 Announce Type: new Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between…
SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
arXiv:2608.07876v1 Announce Type: new Abstract: Autonomous laparoscopic camera control requires continuous understanding of the surgeon’s operative intent…
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization…
Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge
arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided…
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
arXiv:2608.07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer…
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment…
