arXiv:2608.23568v1 Announce Type: new Abstract: Memory and RAG evaluations often treat the answering model’s input as an implementation detail, even…
Tag: cs.AI updates on arXiv.org
LLM Agents Perform Controlled Experiments Using Simulation Models
arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many…
A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
arXiv:2608.23626v1 Announce Type: new Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived…
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
arXiv:2608.06165v3 Announce Type: replace-cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to…
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
arXiv:2607.24555v2 Announce Type: replace-cross Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache, which…
AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization
arXiv:2608.07557v2 Announce Type: replace-cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive…
MOSAIC: Masked Outsourcing of Secure AI Computations
arXiv:2607.29221v2 Announce Type: replace-cross Abstract: We address the challenge of securely and efficiently outsourcing AI computations from a trusted…
TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction
arXiv:2608.01400v2 Announce Type: replace-cross Abstract: Tabular foundation models, driven by in-context learning, have rapidly grown in quality and…
EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology
arXiv:2607.19261v4 Announce Type: replace-cross Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions,…
GraphVid: Interactive Graph-Controllable Video Generation
arXiv:2607.21580v2 Announce Type: replace-cross Abstract: Controllable video generation remains challenging due to the difficulty of specifying precise…
