arXiv:2608.24979v2 Announce Type: replace Abstract: Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most…
Category: cs.AI updates on arXiv.org
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
arXiv:2609.01315v2 Announce Type: replace Abstract: Building an omni-modal foundation model means evaluating it across text, image, video, and audio.…
From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale
arXiv:2609.05758v2 Announce Type: replace Abstract: Conversational assistants can blend retrieval, action selection, escalation, and wording in a single…
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
arXiv:2609.04298v3 Announce Type: replace Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often…
LiFTER: A Grounded Neuro-Symbolic Microscope for Continuous-Time Dynamic Graph Forecasting
arXiv:2608.06765v2 Announce Type: replace Abstract: Continuous-time dynamic graph models predict future links by compressing past interactions into neural…
Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
arXiv:2608.16578v2 Announce Type: replace Abstract: AI agents increasingly operate as part of interacting systems rather than in isolation. As agents…
A Human Audit of OpenAIs AI-Generated Mathematical Proofs
arXiv:2608.14673v3 Announce Type: replace Abstract: We assess 18 chapter-specific reviews of the ten mathematical results announced by OpenAI on 1 August…
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning
arXiv:2608.06411v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language…
Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation
arXiv:2608.15877v3 Announce Type: replace Abstract: Search and recommendation serve a shared discovery objective but encode intent differently. We study…
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
arXiv:2608.05833v3 Announce Type: replace Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph…
