arXiv:2609.22329v1 Announce Type: cross Abstract: Black-Box Optimization (BBO) is often applied in several engineering fields and can utilize an…
Tag: AI
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, 2 new text-to-speech models available now through the Gemini API and Google AI Studio. Flash…
Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch
arXiv:2609.22359v1 Announce Type: cross Abstract: An aligned model asked to hold its answer against a manipulative source must still update on a reliable…
The AI Hype Index: AI loves cheating
Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test.…
AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation
arXiv:2609.22332v2 Announce Type: cross Abstract: Generalizable robot manipulation requires predicting how a scene will evolve, identifying where…
GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents
arXiv:2609.22308v1 Announce Type: cross Abstract: Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in…
Authority-Preserving Evaluation of Medical Vision-Language Assistants
arXiv:2609.22302v1 Announce Type: cross Abstract: Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local…
Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models
arXiv:2609.22293v1 Announce Type: cross Abstract: Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in…
Visual Graph Reasoning via Knowledge Compilation
arXiv:2609.22327v1 Announce Type: cross Abstract: Visual graph reasoning requires answering graph-theoretic questions directly from graph images, where…
Introducing MentalHealthBench
MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
