arXiv:2609.07925v4 Announce Type: replace Abstract: We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently…
Do Not Restart: Residual Completion for Stateful Agent Handoffs
arXiv:2609.13800v2 Announce Type: replace Abstract: Routing and cascades reduce tool-agent cost by transferring control across models, but stateful…
AI News Brief Hourly Summary 2026-09-18 04h : 13 posts
13 posts published in the last hour 01:32HyQuant: Hybrid-Precision Quantization for LLM Attention 01:32EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses 01:32A visual large language foundational model for medical image recognition using clinician-contributed online resources 01:32Iris: Climbing to the Search Frontier…
HyQuant: Hybrid-Precision Quantization for LLM Attention
arXiv:2608.27875v3 Announce Type: replace Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve…
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution…
A visual large language foundational model for medical image recognition using clinician-contributed online resources
arXiv:2609.06914v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing…
Iris: Climbing to the Search Frontier
arXiv:2609.04304v2 Announce Type: replace Abstract: We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales,…
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v2 Announce Type: replace Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture…
Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models
arXiv:2603.19087v3 Announce Type: replace Abstract: Creative ideas often arise by associating remote concepts. Can random associations reliably increase…
Predictive Assistance and the Temporal Dynamics of Exploratory Compression
arXiv:2606.10094v2 Announce Type: replace Abstract: Classical theories of cognition describe problem solving as exploratory search through structured…
Exploratory Responsiveness and Adaptive Rigidity under AI-Assisted Optimization
arXiv:2606.10086v2 Announce Type: replace Abstract: This paper develops a theory of exploratory adaptation under AI-assisted optimization. The central…
Google Deepmind launches interdisciplinary institute to tackle the big questions around AGI
Google Deepmind has founded the Deepmind Institute (DMI), an interdisciplinary research platform focused on AGI. Led by Demis Hassabis, Shane Legg, and…
An Agentic Framework for Neuro-Symbolic Programming
arXiv:2601.00743v2 Announce Type: replace Abstract: Integrating symbolic constraints into deep learning models could make them more robust, interpretable,…
Crusoe raises $3.9B to build massive data centers and small modular ‘AI factories’
The round values the data center giant at $30.9 billion.
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
arXiv:2608.15565v4 Announce Type: replace Abstract: Agents that learn from experience improve at optimization modeling by storing solved trajectories and…
AI News Brief Hourly Summary 2026-09-18 03h : 14 posts
14 posts published in the last hour 00:32LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition 00:32MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use 00:32Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model…
LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
arXiv:2510.08928v2 Announce Type: replace Abstract: Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in…
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
arXiv:2512.24565v5 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of…
