arXiv:2608.14881v1 Announce Type: new Abstract: AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and…
Author: script
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
arXiv:2608.14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation…
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
arXiv:2608.14795v1 Announce Type: new Abstract: An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of…
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
arXiv:2608.14808v1 Announce Type: new Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient,…
Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations
arXiv:2608.14789v1 Announce Type: new Abstract: Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically…
Anthropic increases revenue sevenfold, hits annualized rate above $65 billion
Anthropic’s annualized revenue has topped $65 billion, a sevenfold increase in one year, according to Bloomberg. The company could go public as early as…
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet…
AI News Brief Hourly Summary 2026-08-18 10h : 11 posts
11 posts published in the last hour 07:33Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs 07:33Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation 07:33Advanced modelling and data analytics in aviation 07:33Semantic Uncertainty-Guided…
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs
arXiv:2608.14765v1 Announce Type: new Abstract: Data cleaning without a trusted clean reference is challenging because unusual values may represent either…
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
arXiv:2608.14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current…
