arXiv:2608.22504v1 Announce Type: new Abstract: Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy.…
Tag: cs.AI updates on arXiv.org
When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing
arXiv:2608.22438v1 Announce Type: new Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across…
ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts
arXiv:2608.22510v1 Announce Type: new Abstract: Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue…
Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding
arXiv:2608.22429v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for…
Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA
arXiv:2608.22363v1 Announce Type: new Abstract: Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far…
LLMs for Survey Text Analysis – A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis
arXiv:2608.22417v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet…
WAM-OPD: On-Policy Distillation for World Action Models
arXiv:2608.22364v1 Announce Type: new Abstract: World action models (WAMs) couple visual future prediction with robot action generation, but accelerated…
Where World Models Break: Natural-Input Failure Discovery
arXiv:2608.22421v1 Announce Type: new Abstract: World models predict action-conditioned futures and serve as critical internal simulators for downstream…
HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory
arXiv:2608.22310v1 Announce Type: new Abstract: Long-term memory is crucial for personalized responses and long-horizon agent interactions. Existing…
Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture
arXiv:2608.22347v1 Announce Type: new Abstract: A cognitive architecture is more than the module that reasons: it must also decide how long to think and…
