arXiv:2609.05527v1 Announce Type: new Abstract: Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the…
SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction
arXiv:2609.05511v1 Announce Type: new Abstract: Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most…
AI News Brief Hourly Summary 2026-09-10 07h : 10 posts
10 posts published in the last hour 04:32SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews 04:32RAPID: Reliability-Aware Pair Importance Distillation 04:32ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models 04:32PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of…
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations…
RAPID: Reliability-Aware Pair Importance Distillation
arXiv:2609.05481v1 Announce Type: new Abstract: Inter example relational distillation transfers a teacher’s representation geometry by matching relations…
ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models
arXiv:2609.05461v1 Announce Type: new Abstract: Reward-free latent world models plan by scoring candidate actions with distances in a frozen latent space:…
PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories
arXiv:2609.05488v1 Announce Type: new Abstract: Clinical deterioration unfolds through coupled, partially observed trajectories, not a single diagnostic…
Compiling VGDL into Causal Models
arXiv:2609.05459v1 Announce Type: new Abstract: Reinforcement learning and large language models often struggle to accurately capture the causal mechanics…
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
arXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo,…
CriticGen: Generation-Aware Evaluation as Actionable Feedback
arXiv:2609.05439v1 Announce Type: new Abstract: Current evaluation methods for large language models are coarse-grained and decoupled from generation,…
AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents
arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model…
Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms — teaching models what…
Damage-Aware Bandit Pruning for Vision and Language Transformers
arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose…
LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
LandingAI has shipped Agentic Document Extraction Gen2, a rebuild of its document stack on the DPT-3 model family. Chunks are retired in favor of a…
AI research startup Listen Labs scrubbed a $1.5B funding round for Salesforce talks
Listen Labs walked away from a signed Series C term sheet from Menlo Ventures, sources say.
The AI policy window is open. We need to act.
Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window…
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
Google has open-sourced Mantis, a stack-agnostic toolkit of security review skills for AI coding agents. It runs the full vulnerability lifecycle: sweep…
AI News Brief Hourly Summary 2026-09-10 01h : 5 posts
5 posts published in the last hour 22:32OpenAI adds a prominent AI doomer to its board of directors 22:32Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM 22:32Apple has a new way to prove your iPhone photos aren’t AI slop 22:03GPT-6…
