13 posts published in the last hour 03:32Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches 03:32Disclosure-Gated User Simulation for Companion-Agent Evaluation 03:32Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling…
Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches
arXiv:2609.00946v1 Announce Type: cross Abstract: Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and…
Disclosure-Gated User Simulation for Companion-Agent Evaluation
arXiv:2609.00982v1 Announce Type: cross Abstract: Using a large language model to play the user is now standard in scalable evaluation. It has a…
Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
arXiv:2609.00949v1 Announce Type: cross Abstract: Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public…
From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding
arXiv:2609.00948v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated strong performance in visual question answering with…
The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research
arXiv:2609.00969v1 Announce Type: cross Abstract: We present the zbMATH Open Knowledge Graph, a large-scale RDF knowledge graph (KG) covering more than…
Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting
arXiv:2609.00898v1 Announce Type: cross Abstract: Obtaining labeled data for semantic segmentation in applied settings (e.g., autonomous driving,…
DualStake: Dual-Path Confidence Calibration in Deep Research Agents
arXiv:2609.00935v1 Announce Type: cross Abstract: Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and…
Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
arXiv:2609.00866v1 Announce Type: cross Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational…
Introducing agentic video understanding with Gemini
This post has no text preview — click the link below to read the original article. This article has been indexed from Google DeepMind News Read the original article: Introducing agentic video understanding with Gemini
Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
arXiv:2609.00924v1 Announce Type: cross Abstract: Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and…
Google’s answer to Canva is an AI tool where you prompt instead of design
With Google Pics, Google is pushing deeper into the creative software market dominated by Canva and Adobe, but with a distinctly AI-first approach.
Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO
arXiv:2609.00925v1 Announce Type: cross Abstract: Language models can ignore prompt evidence when it conflicts with memorized knowledge. Post-training can…
AI News Brief Hourly Summary 2026-09-03 05h : 15 posts
15 posts published in the last hour 02:32ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection 02:32Replacing Training with Memory: Listwise Selection for Text-to-SQL 02:32Probabilistic Model Checking of Autoregressive Neural Sequence Models 02:32How AI-native companies turn workflows into operating…
ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection
arXiv:2609.00853v1 Announce Type: cross Abstract: InfRared Small Target Detection (IRSTD) is a challenging task. Relying solely on pixel-level…
Replacing Training with Memory: Listwise Selection for Text-to-SQL
arXiv:2609.00834v1 Announce Type: cross Abstract: Modern Text-to-SQL systems often follow generate-execute-select pipelines, generating multiple candidate…
Probabilistic Model Checking of Autoregressive Neural Sequence Models
arXiv:2609.00838v1 Announce Type: cross Abstract: Test-set accuracy is silent on two issues that matter when deploying autoregressive neural sequence…
How AI-native companies turn workflows into operating capability
Basis, Clay, and Exa Labs use AI agents to improve onboarding, account management, and developer integrations. See what enterprise leaders can apply.
