arXiv:2609.28963v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading…
Tag: cs.AI updates on arXiv.org
From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs
arXiv:2609.28942v1 Announce Type: new Abstract: Personalized value alignment has become increasingly important as large language models (LLMs) are…
When Does Action Credit Need Updating?
arXiv:2609.29007v1 Announce Type: new Abstract: Tool-using agents are continually updated with new interaction data. After each policy update, however,…
MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks
arXiv:2609.29015v1 Announce Type: new Abstract: Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain…
AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining
arXiv:2609.29014v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems can automate alpha factor mining, but their reliance…
PFArena: Benchmarking Language Models for Protein Modification
arXiv:2609.28921v1 Announce Type: new Abstract: Protein modification requires navigating an immense sequence space, yet wet-lab validation remains…
Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise
arXiv:2609.28919v1 Announce Type: new Abstract: Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out…
Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation
arXiv:2609.28859v1 Announce Type: new Abstract: Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and…
RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
arXiv:2609.28850v1 Announce Type: new Abstract: Reproducing a machine learning paper involves most research steps, from installing software and debugging…
Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents
arXiv:2609.28876v1 Announce Type: new Abstract: We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents.…
