arXiv:2608.07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse,…
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses…
TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents
arXiv:2608.07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible…
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return “no evidence” either because a benchmark is clean or because…
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing…
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic…
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be…
GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
arXiv:2608.07881v1 Announce Type: new Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between…
SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
arXiv:2608.07876v1 Announce Type: new Abstract: Autonomous laparoscopic camera control requires continuous understanding of the surgeon’s operative intent…
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization…
AI News Brief Hourly Summary 2026-08-11 10h : 12 posts
12 posts were published in the last hour 7:32 : Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge 7:32 : When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure…
Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge
arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided…
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
arXiv:2608.07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer…
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment…
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it…
How AI is changing the vulnerability response timeline
Artificial intelligence is giving security researchers new ways to examine code, trace unusual behaviour and identify flaws that conventional tools may…
Back to the Future: A workbook time machine for spread sheet creation benchmarks
arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the…
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage…
