arXiv:2608.24192v1 Announce Type: new Abstract: Aligning large language models to human preferences is crucial for real-world deployment but frequently…
Category: cs.AI updates on arXiv.org
Evaluating Multiple LLM Generations with Validated Task Coverage
arXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide several candidate outputs for comparison,…
TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even…
Task-Adaptive Rubrics for GUI Reward Modeling
arXiv:2608.24174v1 Announce Type: new Abstract: Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome…
Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
arXiv:2608.24135v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancing the…
AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL
arXiv:2608.24114v1 Announce Type: new Abstract: Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewards,…
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
arXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and…
Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
arXiv:2608.24188v1 Announce Type: new Abstract: Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context…
ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
arXiv:2608.24103v1 Announce Type: new Abstract: Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two…
EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals
arXiv:2608.24086v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as code agents for scientific and engineering analysis,…
