arXiv:2608.22538v1 Announce Type: new Abstract: Policy-governed agents must interpret case evidence while following an authorized procedure. We present…
Category: cs.AI updates on arXiv.org
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
arXiv:2608.22559v1 Announce Type: new Abstract: Rubrics aim to make language-model evaluation transparent by decomposing response quality into…
Scaling Curriculum Learning For Autonomous Driving
arXiv:2608.22549v1 Announce Type: new Abstract: Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL)…
Small Reasoning Models are Instruction Followers in Function Calling
arXiv:2608.22472v1 Announce Type: new Abstract: Function calling represents the core capability of agentic large language models (LLMs). Existing research…
HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems
arXiv:2608.22512v1 Announce Type: new Abstract: Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations.…
When Does AI for PDEs Yield Scientific Evidence?
arXiv:2608.22504v1 Announce Type: new Abstract: Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy.…
When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing
arXiv:2608.22438v1 Announce Type: new Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across…
ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts
arXiv:2608.22510v1 Announce Type: new Abstract: Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue…
Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding
arXiv:2608.22429v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for…
Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA
arXiv:2608.22363v1 Announce Type: new Abstract: Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far…
