arXiv:2605.29068v2 Announce Type: replace Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in…
Category: cs.AI updates on arXiv.org
BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA
arXiv:2603.28026v4 Announce Type: replace Abstract: Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively…
Role-Aware Artificial Intelligence Across Augmentation and Automation in Human-Machine Symbiosis
arXiv:2605.00440v2 Announce Type: replace Abstract: The evolution of artificial intelligence (AI) has rendered the boundary between humanity and…
Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents
arXiv:2606.08151v3 Announce Type: replace Abstract: Modern large language model (LLM) agents do not simply need longer contexts; they need…
SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision
arXiv:2606.01139v4 Announce Type: replace Abstract: Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints,…
MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution
arXiv:2603.18718v2 Announce Type: replace Abstract: Memory-augmented LLM agents maintain external memory banks to support long-horizon interaction, yet…
MemCoRe: Recovering Evidence from Progressively Compressed Factual Knowledge for Agent Memory
arXiv:2602.07885v3 Announce Type: replace Abstract: Memory systems enable LLM agents to consolidate and retrieve relevant evidence from the factual…
RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
arXiv:2602.05765v3 Announce Type: replace Abstract: Reinforcement learning (RL) has emerged as a critical paradigm for post-training…
OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Heuristic Design
arXiv:2602.13769v4 Announce Type: replace Abstract: Automating heuristic design in complex, experiment-driven domains requires more than iterative…
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
arXiv:2603.08234v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), the safety of LLMs has become a critical…
