arXiv:2609.22308v1 Announce Type: cross Abstract: Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in…
Author: script
Authority-Preserving Evaluation of Medical Vision-Language Assistants
arXiv:2609.22302v1 Announce Type: cross Abstract: Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local…
Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models
arXiv:2609.22293v1 Announce Type: cross Abstract: Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in…
Visual Graph Reasoning via Knowledge Compilation
arXiv:2609.22327v1 Announce Type: cross Abstract: Visual graph reasoning requires answering graph-theoretic questions directly from graph images, where…
Introducing MentalHealthBench
MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
On The Robustness-Resolution Tradeoff In Temporal Quantization Of Event Streams
arXiv:2609.22295v2 Announce Type: cross Abstract: Event pipelines often discretize asynchronous timestamps before learning. This step looks harmless, but…
AI News Brief Hourly Summary 2026-09-23 22h : 18 posts
18 posts published in the last hour 19:33Rethinking Streaming Video Diffusion Model: Context, Execution, and Training 19:32ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI 19:32Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection 19:32Performance vs Consistency: Evaluating a Foundation…
Rethinking Streaming Video Diffusion Model: Context, Execution, and Training
arXiv:2609.22283v1 Announce Type: cross Abstract: Understanding the design space of streaming video diffusion is essential to exploring its potential for…
ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI
arXiv:2609.22285v1 Announce Type: cross Abstract: Adapting language models to new domains via continual pre-training raises a basic evaluation problem: if…
Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection
arXiv:2609.22284v1 Announce Type: cross Abstract: Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two…
