arXiv:2609.22359v1 Announce Type: cross Abstract: An aligned model asked to hold its answer against a manipulative source must still update on a reliable…
Author: script
The AI Hype Index: AI loves cheating
Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test.…
AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation
arXiv:2609.22332v2 Announce Type: cross Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where…
GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents
arXiv:2609.22308v1 Announce Type: cross Abstract: Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in…
Authority-Preserving Evaluation of Medical Vision-Language Assistants
arXiv:2609.22302v1 Announce Type: cross Abstract: Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local…
Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models
arXiv:2609.22293v1 Announce Type: cross Abstract: Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in…
Visual Graph Reasoning via Knowledge Compilation
arXiv:2609.22327v1 Announce Type: cross Abstract: Visual graph reasoning requires answering graph-theoretic questions directly from graph images, where…
Introducing MentalHealthBench
MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
On The Robustness-Resolution Tradeoff In Temporal Quantization Of Event Streams
arXiv:2609.22295v2 Announce Type: cross Abstract: Event pipelines often discretize asynchronous timestamps before learning. This step looks harmless, but…
AI News Brief Hourly Summary 2026-09-23 22h : 18 posts
18 posts published in the last hour 19:33Rethinking Streaming Video Diffusion Model: Context, Execution, and Training 19:32ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI 19:32Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection 19:32Performance vs Consistency: Evaluating a Foundation…
