arXiv:2609.00949v1 Announce Type: cross Abstract: Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public…
Tag: AI
From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding
arXiv:2609.00948v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated strong performance in visual question answering with…
The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research
arXiv:2609.00969v1 Announce Type: cross Abstract: We present the zbMATH Open Knowledge Graph, a large-scale RDF knowledge graph (KG) covering more than…
Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting
arXiv:2609.00898v1 Announce Type: cross Abstract: Obtaining labeled data for semantic segmentation in applied settings (e.g., autonomous driving,…
DualStake: Dual-Path Confidence Calibration in Deep Research Agents
arXiv:2609.00935v1 Announce Type: cross Abstract: Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and…
Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
arXiv:2609.00866v1 Announce Type: cross Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational…
Introducing agentic video understanding with Gemini
This post has no text preview — click the link below to read the original article. This article has been indexed from Google DeepMind News Read the original article: Introducing agentic video understanding with Gemini
Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
arXiv:2609.00924v1 Announce Type: cross Abstract: Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and…
Google’s answer to Canva is an AI tool where you prompt instead of design
With Google Pics, Google is pushing deeper into the creative software market dominated by Canva and Adobe, but with a distinctly AI-first approach.
Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO
arXiv:2609.00925v1 Announce Type: cross Abstract: Language models can ignore prompt evidence when it conflicts with memorized knowledge. Post-training can…
