arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed…
Category: AI
KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs
arXiv:2608.07954v1 Announce Type: new Abstract: Large language models can answer knowledge-intensive questions more reliably when they are grounded with…
Anthropic watermarks all Claude outputs globally with marks that “may persist through some editing”
Anthropic will embed invisible watermarks in all Claude-generated text and sign files using the C2PA standard. New models shipping from August 2026 onward…
Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs
arXiv:2608.07947v1 Announce Type: new Abstract: Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional…
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
arXiv:2608.07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse,…
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses…
TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents
arXiv:2608.07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible…
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return “no evidence” either because a benchmark is clean or because…
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing…
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic…