arXiv:2607.14616v3 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether…
Tag: cs.AI updates on arXiv.org
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
arXiv:2606.16149v3 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer;…
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching
arXiv:2605.29055v2 Announce Type: replace Abstract: This paper describes an approach to hallucination detection and mitigation using a HOPE-inspired…
Towards Human Motion World Models via Executable Behaviour Representations
arXiv:2604.18064v2 Announce Type: replace Abstract: Human motion world models should capture motion’s intentionality by being executable: adaptable to…
Tools as Continuous Flow for Evolving Agentic Reasoning
arXiv:2605.07339v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for…
RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit
arXiv:2606.06027v2 Announce Type: replace Abstract: Community-conditioned language model adaptation needs choices about data collection, community…
Moxia: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning
arXiv:2606.00671v3 Announce Type: replace Abstract: We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining…
Behavior and Representation in Open-Weight Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
arXiv:2512.13374v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) open new perspectives for automation in optimization,…
DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation
arXiv:2507.14267v2 Announce Type: replace Abstract: Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical…
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
arXiv:2509.13450v3 Announce Type: replace Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine…
