arXiv:2605.16024v2 Announce Type: replace Abstract: Desktop GUI agents operate under partial observability: visually similar screens can correspond to…
Tag: cs.AI updates on arXiv.org
Pander Score: A Continuous Measure of Sycophancy as Epistemic Deference
arXiv:2606.07897v2 Announce Type: replace Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user.…
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
arXiv:2604.12616v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual…
Planning under Distribution Shifts with Causal POMDPs
arXiv:2602.23545v3 Announce Type: replace Abstract: In the real world, planning is often challenged by distribution shifts. As such, a model of the…
Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
arXiv:2604.00547v2 Announce Type: replace Abstract: Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a…
Chronos: The AI Co-Historian
arXiv:2604.03553v3 Announce Type: replace Abstract: AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet,…
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
arXiv:2510.08713v3 Announce Type: replace Abstract: Enabling embodied agents to imagine future states is essential for robust and generalizable visual…
TSQueryBench: LLM-as-a-Judge for Time Series Explanations
arXiv:2604.02118v2 Announce Type: replace Abstract: Natural language explanations of time series data are increasingly produced by foundation models in…
MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale
arXiv:2506.01297v5 Announce Type: replace Abstract: Representation learning of geospatial locations remains a core challenge in achieving general…
LLM Enhancement with Domain Expert Mental Model to Reduce LLM Hallucination with Causal Prompt Engineering
arXiv:2509.10818v2 Announce Type: replace Abstract: When consequential decisions depend on knowledge that exists nowhere in writing, LLMs hallucinate not…
