arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and…
Category: AI
Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents
arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic…
Coupled Graph–Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity
arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices…
Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
arXiv:2608.09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source…
KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models
arXiv:2608.09412v1 Announce Type: new Abstract: KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct…
Making Knowledge Distillation Cheap Enough to Run at Scale
This post has no text preview — click the link below to read the original article. This article has been indexed from Hugging Face – Blog Read the original article: Making Knowledge Distillation Cheap Enough to Run at Scale
From Prompt to Harness: Coderlet from Scratch
arXiv:2608.09480v1 Announce Type: new Abstract: A model alone does not determine how a programming agent acts. What the model sees, how actions enter the…
GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models
arXiv:2608.09325v1 Announce Type: new Abstract: Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines…
OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks
arXiv:2608.09380v1 Announce Type: new Abstract: Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools,…
LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling
arXiv:2608.09343v1 Announce Type: new Abstract: Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most…