arXiv:2609.01045v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities as powerful components in agentic…
Category: cs.AI updates on arXiv.org
User Representation via Cross Multi-source Behavior Pre-training for Mobile Games
arXiv:2609.01057v1 Announce Type: new Abstract: User representation pre-training has become a fundamental paradigm for alleviating data sparsity in…
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
arXiv:2609.01058v1 Announce Type: new Abstract: Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold…
WorldBench: Culturally Grounded Benchmark for Multilingual Agents
arXiv:2609.01056v1 Announce Type: new Abstract: Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing…
Few-Shot Out of Domain Intent Detection with Covariance Corrected Mahalanobis Distance
arXiv:2609.00961v1 Announce Type: new Abstract: Conversational agents like chatbots and voice assistants are trained to understand and respond to user…
Figures as Programs: Recursive Generation of Editable Scientific Figures
arXiv:2609.01006v1 Announce Type: new Abstract: Scientific methodology figures are essential for communicating complex methods clearly, yet creating them…
Data-Driven Persona-Conditioned Agents for A/B Test Simulation
arXiv:2609.01038v1 Announce Type: new Abstract: A/B testing is the gold standard for evaluating product changes, but each experiment requires real user…
Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees
arXiv:2609.01035v1 Announce Type: new Abstract: Recursive LLM agents can broaden their search by spawning specialists. Some branches later request tools…
CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins
arXiv:2609.00967v1 Announce Type: new Abstract: As large language models increasingly act through external tools, deciding when to call a tool has become…
RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation
arXiv:2609.00918v1 Announce Type: new Abstract: Large language models are increasingly used as interactive recommender assistants. Their evaluation should…
