arXiv:2608.12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must…
Tag: cs.AI updates on arXiv.org
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
arXiv:2608.11977v1 Announce Type: new Abstract: Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed…
OEIS Open: How many conjectures can language models turn into theorems?
arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized…
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
arXiv:2608.11775v1 Announce Type: new Abstract: Gist-based context compression—summarising older conversation history into compact representations—is…
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
arXiv:2608.11949v1 Announce Type: new Abstract: Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent…
Policy-as-logic for robust reasoning over rules
arXiv:2608.11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance,…
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape…
Proportional Analogies on Probability Distributions via Bayesian Updating
arXiv:2608.11724v1 Announce Type: new Abstract: Analogies are quaternary relations of the form “A is to B as C is to D”. Among the various formalizations…
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
arXiv:2608.11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which…
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
arXiv:2608.11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing…
