arXiv:2506.06793v2 Announce Type: replace-cross Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online…
Category: AI
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
arXiv:2608.06243v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large…
Towards a Theoretical Understanding of Two Tower Recommendation Models
arXiv:2403.00802v2 Announce Type: replace-cross Abstract: Production-grade recommender systems rely heavily on a large-scale corpus used by online media…
Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction
arXiv:2411.00028v3 Announce Type: replace-cross Abstract: Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic…
Contextual Information Policy Optimization for Search Agents
arXiv:2608.06128v2 Announce Type: replace Abstract: Search agents extend large language models beyond static parametric memory by enabling them to acquire…
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
arXiv:2608.06305v2 Announce Type: replace Abstract: Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed…
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
arXiv:2608.03020v2 Announce Type: replace Abstract: Parameter-efficient post-training reduces the number of trainable parameters, but still requires…
Recursive Synthesis for Long-Horizon Terminal Tasks
arXiv:2608.05466v2 Announce Type: replace Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing…
SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
arXiv:2608.05204v2 Announce Type: replace Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata,…
Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
arXiv:2608.03629v2 Announce Type: replace Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an…