arXiv:2609.20004v1 Announce Type: cross Abstract: Reward-based reinforcement learning for language models, exemplified by Group Relative Policy…
Tag: cs.AI updates on arXiv.org
Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
arXiv:2609.19985v1 Announce Type: cross Abstract: Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While…
Efficiently Distributed Federated Learning
arXiv:2609.19972v1 Announce Type: cross Abstract: Federated Learning (FL) is experiencing a substantial research interest, with many frameworks being…
AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models
arXiv:2609.19991v1 Announce Type: cross Abstract: Omni models can describe video content, but can they locate events in time, preserve event order, and…
Zarya: A Hybrid Autoregressive–Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference
arXiv:2609.19868v1 Announce Type: cross Abstract: Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while…
Learning and Transferring Closed-Loop Robot Software
arXiv:2609.19906v1 Announce Type: cross Abstract: Closed-loop robot policies require observation processing, state management, and situation-dependent…
KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms
arXiv:2609.19916v1 Announce Type: cross Abstract: Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language…
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
arXiv:2609.19883v1 Announce Type: cross Abstract: Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific…
ClashBench: Conflicts Leading Agents to Seize and Harm
arXiv:2609.19892v1 Announce Type: cross Abstract: As agent systems become more widely used, multiple agent sessions increasingly run alongside…
Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision
arXiv:2609.19846v1 Announce Type: cross Abstract: As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain…
