arXiv:2609.05374v1 Announce Type: new Abstract: Computer-use agents have advanced on benchmarks like OSWorld and AndroidWorld, but still act mostly…
Category: cs.AI updates on arXiv.org
Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
arXiv:2609.05395v1 Announce Type: new Abstract: Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise…
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
arXiv:2609.05385v1 Announce Type: new Abstract: LLM decision components that can operate within agent workflows often produce action-relevant…
Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education
arXiv:2609.05346v1 Announce Type: new Abstract: The integration of GenAI tools into higher education assessment raises important questions about how…
LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams
arXiv:2609.05327v1 Announce Type: new Abstract: Quantum circuits are central to implementing quantum algorithms on quantum devices, where quantum gates…
Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models
arXiv:2609.05333v1 Announce Type: new Abstract: A transformer language model assigns a single, context-independent vector to a word type at its embedding…
Does Your Agent’s Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
arXiv:2609.05339v1 Announce Type: new Abstract: Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still…
Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness
arXiv:2609.05314v1 Announce Type: new Abstract: Building automation systems generate rich sensor data yet remain insight-poor because heterogeneous point…
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
arXiv:2609.05295v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but…
GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity
arXiv:2609.05284v1 Announce Type: new Abstract: Recent years have witnessed great advances in the reasoning ability of Large Language Models (LLMs).…
