arXiv:2608.14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing…
MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
arXiv:2608.14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based…
JarvisBench: Always-on Intelligence Between Humans and Agents
arXiv:2608.14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This…
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
arXiv:2608.14841v1 Announce Type: new Abstract: Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text,…
AI systems quietly drop user instructions when they compress context
When AI systems condense long conversations, they drop an average of 83 percent of user rules, like “don’t send emails without my approval.” Penn State…
Personalized Auto-Research: Towards a True AI Co-Scientist
arXiv:2608.14881v1 Announce Type: new Abstract: AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and…
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
arXiv:2608.14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation…
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
arXiv:2608.14795v1 Announce Type: new Abstract: An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of…
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
arXiv:2608.14808v1 Announce Type: new Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient,…
Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations
arXiv:2608.14789v1 Announce Type: new Abstract: Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically…
Anthropic increases revenue sevenfold, hits annualized rate above $65 billion
Anthropic’s annualized revenue has topped $65 billion, a sevenfold increase in one year, according to Bloomberg. The company could go public as early as…
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet…
AI News Brief Hourly Summary 2026-08-18 10h : 11 posts
11 posts published in the last hour 07:33Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs 07:33Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation 07:33Advanced modelling and data analytics in aviation 07:33Semantic Uncertainty-Guided…
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs
arXiv:2608.14765v1 Announce Type: new Abstract: Data cleaning without a trusted clean reference is challenging because unusual values may represent either…
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
arXiv:2608.14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current…
Advanced modelling and data analytics in aviation
arXiv:2608.14746v1 Announce Type: new Abstract: The aviation industry characterized by its stringent safety standards has seen a growing need for…
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
arXiv:2608.14707v1 Announce Type: new Abstract: As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents…
Haut.AI Expands Across 4,000 O Boticário Stores After Pilot Lifts Skincare Order Value 80%
Estonian skin-analysis vendor Haut.AI is moving from a 24-store experiment to a national retail deployment. On August 18, 2026, the company and Brazil’s…
