arXiv:2608.14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This…
Category: cs.AI updates on arXiv.org
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
arXiv:2608.14841v1 Announce Type: new Abstract: Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text,…
Personalized Auto-Research: Towards a True AI Co-Scientist
arXiv:2608.14881v1 Announce Type: new Abstract: AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and…
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
arXiv:2608.14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation…
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
arXiv:2608.14795v1 Announce Type: new Abstract: An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of…
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
arXiv:2608.14808v1 Announce Type: new Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient,…
Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations
arXiv:2608.14789v1 Announce Type: new Abstract: Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically…
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet…
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs
arXiv:2608.14765v1 Announce Type: new Abstract: Data cleaning without a trusted clean reference is challenging because unusual values may represent either…
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
arXiv:2608.14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current…
