arXiv:2609.00805v1 Announce Type: new Abstract: LLM-based software systems increasingly require effective “evals” as quality gates in the development…
Category: cs.AI updates on arXiv.org
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
arXiv:2609.00787v1 Announce Type: new Abstract: Humans need to study only a handful of well-written textbooks to master a discipline and attempt its…
One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning
arXiv:2609.00813v1 Announce Type: new Abstract: While reinforcement learning has enabled LLM-based search agents to invoke external tools, existing…
DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
arXiv:2609.00768v1 Announce Type: new Abstract: Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver…
Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs
arXiv:2609.00738v1 Announce Type: new Abstract: Inference-time search with large language models (LLMs) often concentrates on a small set of structurally…
S^3martCirc: Self-supervised Smart Circuit Discovery
arXiv:2609.00755v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks, from text…
Agentic Empirical Asset Pricing: Methodological Foundations
arXiv:2609.00731v1 Announce Type: new Abstract: Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical…
ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents
arXiv:2609.00749v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents require context assembly: the runtime must decide what to…
Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks
arXiv:2609.00763v1 Announce Type: new Abstract: Hierarchical Knowledge graph (KG)-based retrieval augmented generation (RAG) has emerged as a powerful…
SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification
arXiv:2609.00728v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown remarkable promise in translating and reformulating complex…
