arXiv:2609.00046v1 Announce Type: cross Abstract: Rapid and reliable disaster mapping of impacted areas, damaged infrastructure, and affected populations…
Category: cs.AI updates on arXiv.org
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
arXiv:2609.00048v1 Announce Type: cross Abstract: GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use…
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
arXiv:2609.00014v1 Announce Type: cross Abstract: Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However,…
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
arXiv:2609.00038v1 Announce Type: cross Abstract: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final…
EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation
arXiv:2609.01526v1 Announce Type: new Abstract: Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM…
When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
arXiv:2609.01519v1 Announce Type: new Abstract: Interactive simulations increasingly evaluate policies in markets populated by language-model agents.…
Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
arXiv:2609.01552v1 Announce Type: new Abstract: Scientific equation discovery has long been central to scientific progress, proceeding through iterative…
InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information
arXiv:2608.29632v1 Announce Type: cross Abstract: Competitive programming is increasingly being used to evaluate the algorithmic reasoning capabilities of…
Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers
arXiv:2609.01567v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them…
Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations
arXiv:2609.01408v1 Announce Type: new Abstract: A fundamental challenge in artificial intelligence is the transformation of observations into explicit…
