arXiv:2608.13463v2 Announce Type: replace-cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often…
X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
arXiv:2608.16658v2 Announce Type: replace-cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their…
Pre-training Visual Dexterity in Simulation
arXiv:2608.15917v2 Announce Type: replace-cross Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this…
Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules
arXiv:2608.22642v2 Announce Type: replace-cross Abstract: Despite recent advances in molecular foundation models, several limitations remain, such as…
Anthropic wants to do for physical hardware what its Model Context Protocol did for software
Anthropic’s Model Hardware Standard (MHS) gives AI agents a unified interface to physical devices like robotic arms and lab instruments. In early tests,…
Complexity Induction: Compositional Generalization via Structured Training Distortion
arXiv:2608.21464v2 Announce Type: replace-cross Abstract: We demonstrate that structured distortion of training data – which we term complexity induction…
A 12-CNOT Double Qubit Excitation Gate
arXiv:2608.11733v3 Announce Type: replace-cross Abstract: Effective implementation of high-level quantum gates is essential for practical quantum…
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
arXiv:2608.11698v3 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level…
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
arXiv:2608.07945v2 Announce Type: replace-cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage…
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
arXiv:2608.10954v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign…
Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown
Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables,…
M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
arXiv:2608.12196v2 Announce Type: replace-cross Abstract: Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet…
When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
arXiv:2608.04893v2 Announce Type: replace-cross Abstract: Multi-agent LLM systems relay key-value caches instead of text and credit their gains to…
REPREC: Representation Driven Parameter-Efficient Recommendation System
arXiv:2607.24845v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have been applied to sequential recommendation by formulating it as…
TriShieldRAG: 3 Rings, One Blind Spot in Layered Defenses for Retrieval-Augmented Generation
arXiv:2607.23838v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) grounds LLM answers in query-time retrieved documents, so…
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
arXiv:2607.27747v2 Announce Type: replace-cross Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive…
OpenAI cuts off Cursor after SpaceX acquisition, citing Musk’s history of breaking contracts
OpenAI is cutting off the AI coding tool Cursor after SpaceX acquired the company, citing Elon Musk’s track record of breaking contracts. Cursor…
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
arXiv:2608.05745v2 Announce Type: replace-cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while…
