arXiv:2609.18991v1 Announce Type: new Abstract: Multimodal LLMs report strong performance on scientific reasoning benchmarks, yet most treat perception…
Category: cs.AI updates on arXiv.org
Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows
arXiv:2609.18820v1 Announce Type: new Abstract: Agentic workflows now make consequential decisions in regulated settings, and the governance placed around…
Suppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing
arXiv:2609.18985v1 Announce Type: new Abstract: Knowledge-editing benchmarks certify local correctness, whether an edited model produces the new fact on…
Function Lives Where Variance Doesn’t: Task-Weighted Charts of a Language Model’s Computation
arXiv:2609.18989v1 Announce Type: new Abstract: How many dimensions does a language model’s computation actually use? The question is ill-posed until one…
Which LLM is Best for Translating Natural Language Goals to PDDL
arXiv:2609.18731v1 Announce Type: new Abstract: Bridging the gap between human intent and machine execution remains a challenge in automated planning,…
Clueing up LLMs with Tool-Augmented Deductive Reasoning
arXiv:2609.18736v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), performing logically consistent deductive…
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
arXiv:2609.18779v1 Announce Type: new Abstract: Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate…
Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale
arXiv:2609.18769v1 Announce Type: new Abstract: Correctly answering a question grounded in normative documents often depends on information outside any…
Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
arXiv:2609.18723v1 Announce Type: new Abstract: We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding…
Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making
arXiv:2609.18591v1 Announce Type: new Abstract: In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction…
