arXiv:2608.16370v1 Announce Type: new Abstract: Task completion is the standard metric for evaluating context compression, yet it is incomplete:…
Category: cs.AI updates on arXiv.org
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
arXiv:2608.16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but…
DriveCache: Action-Aware Caching for Driving World Model Inference
arXiv:2608.16354v1 Announce Type: new Abstract: Driving video generation models support autonomous-driving development by predicting controllable future…
BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics
arXiv:2608.16211v1 Announce Type: new Abstract: Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research…
Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior
arXiv:2608.16196v1 Announce Type: new Abstract: Personalized game generation requires inferring a player’s abilities and behavioral style from how they…
Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain
arXiv:2608.16164v1 Announce Type: new Abstract: Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early…
Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication
arXiv:2608.16192v1 Announce Type: new Abstract: Generative visual-token communication reduces transmission load by sending only selected discrete tokens…
Competing at Every Price Point with Agentic Evolution over a Menu of LLMs
arXiv:2608.16207v1 Announce Type: new Abstract: Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior…
Assessing LLMs’ mathematical abilities requires understanding the various mechanisms of mathematical creativity
arXiv:2608.16118v1 Announce Type: new Abstract: How should we assess whether large language models can perform mathematical invention? I argue that this…
FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
arXiv:2608.16148v1 Announce Type: new Abstract: Multi-view multi-label feature selection aims to identify a compact and informative feature subset from…
