arXiv:2608.14943v1 Announce Type: new Abstract: Agent skills are often injected in full on every request, increasing token cost. We compare four…
Author: script
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks
arXiv:2608.14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but…
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
arXiv:2608.14940v1 Announce Type: new Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count…
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
arXiv:2608.14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in…
LG Hosts NVIDIA at Seoul Robot Data Factory as 100,000-Hour Training Push Takes Shape
LG Electronics hosted senior NVIDIA officials at its new robot Data Factory in Seoul on August 18, 2026, announcing an accelerated robotics collaboration…
Small Models Scout Bottleneck Order for Large-Model Data Control
arXiv:2608.14936v1 Announce Type: new Abstract: Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether…
AI News Brief Hourly Summary 2026-08-18 11h : 13 posts
13 posts published in the last hour 08:32Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning 08:32MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment 08:32JarvisBench: Always-on Intelligence Between Humans and Agents 08:32What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal…
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
arXiv:2608.14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing…
MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
arXiv:2608.14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based…
JarvisBench: Always-on Intelligence Between Humans and Agents
arXiv:2608.14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This…
