arXiv:2609.01198v2 Announce Type: replace Abstract: Repeated banking interactions require assistants to maintain complete, current, and traceable customer…
Category: cs.AI updates on arXiv.org
OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
arXiv:2609.00015v2 Announce Type: replace Abstract: AI agents powered by large language models are evolving from isolated assistants into heterogeneous…
Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs
arXiv:2609.00575v2 Announce Type: replace Abstract: Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand…
Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate
arXiv:2609.01168v2 Announce Type: replace Abstract: Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW)…
When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems
arXiv:2608.27984v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) have recently moved from static workflows toward dynamically generated…
Can escalation channels redirect reward hacking toward defect disclosure?
arXiv:2608.29460v2 Announce Type: replace Abstract: When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or…
Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling
arXiv:2608.29291v3 Announce Type: replace Abstract: Unified multimodal models jointly support understanding and generation, but incur substantial…
Automated Researchers Can Mitigate Well-characterized Alignment Failures
arXiv:2608.28945v3 Announce Type: replace Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard…
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
arXiv:2608.27463v2 Announce Type: replace Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’…
FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
arXiv:2608.20574v3 Announce Type: replace Abstract: We introduce FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned…
