arXiv:2608.14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but…
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
arXiv:2608.14940v1 Announce Type: new Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count…
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
arXiv:2608.14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in…
LG Hosts NVIDIA at Seoul Robot Data Factory as 100,000-Hour Training Push Takes Shape
LG Electronics hosted senior NVIDIA officials at its new robot Data Factory in Seoul on August 18, 2026, announcing an accelerated robotics collaboration…
Small Models Scout Bottleneck Order for Large-Model Data Control
arXiv:2608.14936v1 Announce Type: new Abstract: Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether…
AI News Brief Hourly Summary 2026-08-18 11h : 13 posts
13 posts published in the last hour 08:32Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning 08:32MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment 08:32JarvisBench: Always-on Intelligence Between Humans and Agents 08:32What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal…
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
arXiv:2608.14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing…
MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
arXiv:2608.14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based…
JarvisBench: Always-on Intelligence Between Humans and Agents
arXiv:2608.14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This…
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
arXiv:2608.14841v1 Announce Type: new Abstract: Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text,…
AI systems quietly drop user instructions when they compress context
When AI systems condense long conversations, they drop an average of 83 percent of user rules, like “don’t send emails without my approval.” Penn State…
Personalized Auto-Research: Towards a True AI Co-Scientist
arXiv:2608.14881v1 Announce Type: new Abstract: AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and…
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
arXiv:2608.14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation…
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
arXiv:2608.14795v1 Announce Type: new Abstract: An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of…
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
arXiv:2608.14808v1 Announce Type: new Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient,…
Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations
arXiv:2608.14789v1 Announce Type: new Abstract: Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically…
Anthropic increases revenue sevenfold, hits annualized rate above $65 billion
Anthropic’s annualized revenue has topped $65 billion, a sevenfold increase in one year, according to Bloomberg. The company could go public as early as…
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet…
