arXiv:2608.19817v1 Announce Type: cross Abstract: Conventional convolutional kernels are typically defined on fixed discrete grids, limiting their ability…
Category: cs.AI updates on arXiv.org
MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents
arXiv:2608.19803v1 Announce Type: cross Abstract: Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often…
Repo0: Design-Driven Zero-to-All Code Generation
arXiv:2608.19854v1 Announce Type: cross Abstract: Large language model agents have made substantial progress in code generation, yet most existing systems…
An Irreducible Quantum Advantage in Aligning World Models with Reality
arXiv:2608.19779v1 Announce Type: cross Abstract: World models provide digital simulacra of the true world, allowing agents to be trained and tested…
Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW
arXiv:2608.19762v1 Announce Type: cross Abstract: A minibatch can influence training beyond the update at which it is observed because AdamW stores past…
Distilling Aggregate Mobility Statistics into a Language Model Policy for Post-Event Crowd Simulation
arXiv:2608.19778v1 Announce Type: cross Abstract: Pedestrian simulators need a behaviour rule for every agent, but privacy usually limits the data for…
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
arXiv:2608.19776v1 Announce Type: cross Abstract: Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an…
LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
arXiv:2608.19800v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive…
Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
arXiv:2608.19739v1 Announce Type: cross Abstract: Multimodal LLMs can see a document, but they often can’t read it reliably. Small text, tables, visual…
Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
arXiv:2608.19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld),…
