arXiv:2608.17687v1 Announce Type: new Abstract: Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the…
Category: cs.AI updates on arXiv.org
Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents
arXiv:2608.17718v1 Announce Type: new Abstract: Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting,…
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch
arXiv:2608.17684v1 Announce Type: new Abstract: Self-evolving agents turn experience into reusable skills, workflows, or memories, but post-evolution…
GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities
arXiv:2608.17665v1 Announce Type: new Abstract: LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such…
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing
arXiv:2608.17638v1 Announce Type: new Abstract: What a reasoning model writes is only a partial record of the process that produces it. We introduce a…
Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models
arXiv:2608.17634v1 Announce Type: new Abstract: The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and…
LLM-Derived Preference Judgments Are Not Self-Consistent
arXiv:2608.17644v1 Announce Type: new Abstract: Agents increasingly interpret a person’s natural-language preferences by querying an LLM for numerical…
Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)
arXiv:2608.17625v1 Announce Type: new Abstract: Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale.…
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making
arXiv:2608.17574v1 Announce Type: new Abstract: How cautious should an agent be while it is still learning its environment? We propose RATTL…
When to Review: Spaced Repetition for Continual Pre-Training of Language Models
arXiv:2608.17530v1 Announce Type: new Abstract: Continual pre-training of large language models must acquire new information without erasing old…
