arXiv:2608.13463v2 Announce Type: replace-cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often…
Category: AI
X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
arXiv:2608.16658v2 Announce Type: replace-cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their…
Pre-training Visual Dexterity in Simulation
arXiv:2608.15917v2 Announce Type: replace-cross Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this…
Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules
arXiv:2608.22642v2 Announce Type: replace-cross Abstract: Despite recent advances in molecular foundation models, several limitations remain, such as…
Anthropic wants to do for physical hardware what its Model Context Protocol did for software
Anthropic’s Model Hardware Standard (MHS) gives AI agents a unified interface to physical devices like robotic arms and lab instruments. In early tests,…
Complexity Induction: Compositional Generalization via Structured Training Distortion
arXiv:2608.21464v2 Announce Type: replace-cross Abstract: We demonstrate that structured distortion of training data – which we term complexity induction…
A 12-CNOT Double Qubit Excitation Gate
arXiv:2608.11733v3 Announce Type: replace-cross Abstract: Effective implementation of high-level quantum gates is essential for practical quantum…
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
arXiv:2608.11698v3 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level…
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
arXiv:2608.07945v2 Announce Type: replace-cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage…
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
arXiv:2608.10954v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign…
