arXiv:2608.23497v1 Announce Type: new Abstract: Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content,…
Category: cs.AI updates on arXiv.org
Characterizing Necessary Losers to Explain Tournaments Losers
arXiv:2608.23446v1 Announce Type: new Abstract: We study the problem of formally explaining why a candidate was not selected by a given tournament rule,…
StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
arXiv:2608.23475v1 Announce Type: new Abstract: As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot…
Walking on the DARKSIDE
arXiv:2608.23370v1 Announce Type: new Abstract: Large Language Models (LLMs) recognise patterns but do not natively track the path of exclusions that a…
Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting
arXiv:2608.23373v1 Announce Type: new Abstract: Forecasting long-range influenza-like illness (ILI) matters for public health readiness. Publicly…
Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning
arXiv:2608.23318v1 Announce Type: new Abstract: Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a…
MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction
arXiv:2608.23397v1 Announce Type: new Abstract: Interactive clinical agents must gather decisive evidence and convert it into grounded actions under…
SkillAlchemy: Open-World Agent Skill Creation
arXiv:2608.23417v1 Announce Type: new Abstract: Agent skills are reusable procedural artifacts that extend language agents with specialized workflows,…
EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models
arXiv:2608.23313v1 Announce Type: new Abstract: Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses,…
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
arXiv:2608.23264v1 Announce Type: new Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these…
