arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model’s native hidden space, requiring probes, sparse…
Tag: cs.AI updates on arXiv.org
Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
arXiv:2608.09512v1 Announce Type: new Abstract: Active inference offers a unified framework for perception, learning, and action, but scaling discrete…
verdi: retrieval is not transfer for continual world model optimization
arXiv:2608.09537v1 Announce Type: new Abstract: Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence.…
Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents
arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable…
The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games
arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and…
Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents
arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic…
Coupled Graph–Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity
arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices…
Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
arXiv:2608.09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source…
KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models
arXiv:2608.09412v1 Announce Type: new Abstract: KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct…
From Prompt to Harness: Coderlet from Scratch
arXiv:2608.09480v1 Announce Type: new Abstract: A model alone does not determine how a programming agent acts. What the model sees, how actions enter the…