arXiv:2608.23258v1 Announce Type: cross Abstract: We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within a…
Category: cs.AI updates on arXiv.org
A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
arXiv:2608.24825v1 Announce Type: new Abstract: The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have…
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
arXiv:2608.24876v1 Announce Type: new Abstract: Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the…
Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA
arXiv:2608.24824v1 Announce Type: new Abstract: Large language models are increasingly used for knowledge graph question answering (KGQA), but can fail to…
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
arXiv:2608.24804v1 Announce Type: new Abstract: We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model…
Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav
arXiv:2608.24764v1 Announce Type: new Abstract: Large language model agents are moving beyond conventional retrieval-augmented generation toward direct…
RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
arXiv:2608.24758v1 Announce Type: new Abstract: Discovering stable neuron behavior across entire domains remains a challenge in mechanistic…
CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
arXiv:2608.24794v1 Announce Type: new Abstract: Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither…
Meta$^n$: Recursive Self-Improvement through Emergent Depth
arXiv:2608.24735v1 Announce Type: new Abstract: Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a…
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
arXiv:2608.24777v1 Announce Type: new Abstract: LLM-based agents can interact with external environments through tool invocation, but this capability also…
