arXiv:2609.27277v1 Announce Type: new Abstract: Time series agents answer analytical questions by calling external tools, and which tools they carry is…
Tag: cs.AI updates on arXiv.org
Enhancing Small Language Models for Power Outage Report Generation via Minimum Risk Training
arXiv:2609.27197v1 Announce Type: new Abstract: Minimum Risk Training (MRT) enables neural machine translation models to directly optimize sequence-level…
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
arXiv:2609.27273v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments…
DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents
arXiv:2609.27276v1 Announce Type: new Abstract: Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose…
XLOG: A CUDA-Native Engine for Neurosymbolic Integration
arXiv:2609.27203v1 Announce Type: new Abstract: xlog is a CUDA-native logic programming engine integrating neural perception with deterministic Datalog,…
Math Reasoning in LLMs is Organized by Approach, Not Topic
arXiv:2609.27041v1 Announce Type: new Abstract: Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their…
Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit
arXiv:2609.27087v1 Announce Type: new Abstract: Organizations increasingly use LLMs for policy, compliance, risk, and operational decision support,…
Provably Complete Generalized Planning with LLMs
arXiv:2609.27105v1 Announce Type: new Abstract: Generalized planning aims to compute a plan that solves all instances of a planning domain. Recent work…
Propose, Don’t Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
arXiv:2609.27051v1 Announce Type: new Abstract: Language-model agents now run the whole of quantitative factor research: they propose investment factors,…
Do We Need Complex Topology Control? Distinct-Peer Random Routing Improves Cost-Efficiency in Sparse Multi-Agent Debate
arXiv:2609.27150v1 Announce Type: new Abstract: Multi-agent debate (MAD) has emerged as a promising paradigm for improving the reasoning accuracy of large…
