arXiv:2608.07440v1 Announce Type: new Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a…
Category: cs.AI updates on arXiv.org
Multimodal Drivers’ Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems
arXiv:2608.06378v1 Announce Type: cross Abstract: Driver emotions can affect risk perception, decision-making, and vehicle control under complex road…
Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation
arXiv:2608.05726v1 Announce Type: cross Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge,…
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
arXiv:2608.07449v1 Announce Type: new Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These…
Interaction Creates Dynamical AI Behavior Absent in Isolation
arXiv:2608.07457v1 Announce Type: new Abstract: What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We…
TEPA: Revoking Stale Memories for Conflict-Robust Language Agents
arXiv:2608.07429v1 Announce Type: new Abstract: Long-term memory enables language agents to reuse past facts, preferences, and task experience.…
PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
arXiv:2608.07438v1 Announce Type: new Abstract: Human-like cognition does not select past experience by topical similarity alone: affective significance…
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
arXiv:2608.07436v1 Announce Type: new Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular…
A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy
arXiv:2608.07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count—a…
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model…