arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution…
Category: cs.AI updates on arXiv.org
A visual large language foundational model for medical image recognition using clinician-contributed online resources
arXiv:2609.06914v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing…
Iris: Climbing to the Search Frontier
arXiv:2609.04304v2 Announce Type: replace Abstract: We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales,…
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v2 Announce Type: replace Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture…
Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models
arXiv:2603.19087v3 Announce Type: replace Abstract: Creative ideas often arise by associating remote concepts. Can random associations reliably increase…
Predictive Assistance and the Temporal Dynamics of Exploratory Compression
arXiv:2606.10094v2 Announce Type: replace Abstract: Classical theories of cognition describe problem solving as exploratory search through structured…
Exploratory Responsiveness and Adaptive Rigidity under AI-Assisted Optimization
arXiv:2606.10086v2 Announce Type: replace Abstract: This paper develops a theory of exploratory adaptation under AI-assisted optimization. The central…
An Agentic Framework for Neuro-Symbolic Programming
arXiv:2601.00743v2 Announce Type: replace Abstract: Integrating symbolic constraints into deep learning models could make them more robust, interpretable,…
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
arXiv:2608.15565v4 Announce Type: replace Abstract: Agents that learn from experience improve at optimization modeling by storing solved trajectories and…
LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
arXiv:2510.08928v2 Announce Type: replace Abstract: Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in…
