arXiv:2609.00982v1 Announce Type: cross Abstract: Using a large language model to play the user is now standard in scalable evaluation. It has a…
Tag: cs.AI updates on arXiv.org
Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
arXiv:2609.00949v1 Announce Type: cross Abstract: Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public…
From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding
arXiv:2609.00948v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated strong performance in visual question answering with…
The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research
arXiv:2609.00969v1 Announce Type: cross Abstract: We present the zbMATH Open Knowledge Graph, a large-scale RDF knowledge graph (KG) covering more than…
Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting
arXiv:2609.00898v1 Announce Type: cross Abstract: Obtaining labeled data for semantic segmentation in applied settings (e.g., autonomous driving,…
DualStake: Dual-Path Confidence Calibration in Deep Research Agents
arXiv:2609.00935v1 Announce Type: cross Abstract: Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and…
Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
arXiv:2609.00866v1 Announce Type: cross Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational…
Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
arXiv:2609.00924v1 Announce Type: cross Abstract: Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and…
Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO
arXiv:2609.00925v1 Announce Type: cross Abstract: Language models can ignore prompt evidence when it conflicts with memorized knowledge. Post-training can…
ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection
arXiv:2609.00853v1 Announce Type: cross Abstract: InfRared Small Target Detection (IRSTD) is a challenging task. Relying solely on pixel-level…
