arXiv:2608.07954v1 Announce Type: new Abstract: Large language models can answer knowledge-intensive questions more reliably when they are grounded with…
Tag: AI
Anthropic watermarks all Claude outputs globally with marks that “may persist through some editing”
Anthropic will embed invisible watermarks in all Claude-generated text and sign files using the C2PA standard. New models shipping from August 2026 onward…
Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs
arXiv:2608.07947v1 Announce Type: new Abstract: Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional…
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
arXiv:2608.07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse,…
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses…
TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents
arXiv:2608.07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible…
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return “no evidence” either because a benchmark is clean or because…
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing…
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic…
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be…
