arXiv:2608.08240v1 Announce Type: new Abstract: This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI…
Category: cs.AI updates on arXiv.org
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment
arXiv:2608.08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter…
A Minimal $\kappa$–$\tau$ Logic for Risk-Sensitive Abduction
arXiv:2608.08192v1 Announce Type: new Abstract: Standard approaches to abductive reasoning can retain multiple candidate explanations, but they do not…
Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets
arXiv:2608.08189v1 Announce Type: new Abstract: LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks…
Quantization Degradation in Large Language Models: A Signal-Noise Perspective
arXiv:2608.08188v1 Announce Type: new Abstract: Post-training quantization reduces the deployment cost of large language models, yet how severely a…
Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models
arXiv:2608.08199v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and…
Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges
arXiv:2608.08184v1 Announce Type: new Abstract: Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but…
When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs
arXiv:2608.08159v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive…
A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning
arXiv:2608.08158v1 Announce Type: new Abstract: Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement…
Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation
arXiv:2608.08146v1 Announce Type: new Abstract: The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long…
