arXiv:2608.03952v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to provide conversational practice for…
Tag: cs.AI updates on arXiv.org
Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning
arXiv:2608.23982v2 Announce Type: replace Abstract: Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it…
HANIA: Planner-Guided Multimodal Graph Evidence Selection for Grounded Question Answering
arXiv:2608.29088v2 Announce Type: replace Abstract: Multimodal question answering remains sensitive to noisy, incomplete, and weakly grounded evidence.…
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
arXiv:2609.09418v3 Announce Type: replace Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing…
Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
arXiv:2607.11433v2 Announce Type: replace Abstract: Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer…
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
arXiv:2606.27161v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but…
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research
arXiv:2605.26114v3 Announce Type: replace Abstract: We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday…
Warranted Attention: Learning What to Pass from Attention to Prediction
arXiv:2606.30139v4 Announce Type: replace Abstract: Relevance of information read by attention does not guarantee that its contribution benefits the…
When do prophets profit in prediction markets?
arXiv:2607.06166v3 Announce Type: replace Abstract: Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of…
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks
arXiv:2607.02175v2 Announce Type: replace Abstract: Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations…
