Zhipu’s release note for GLM-5.3 contains a sentence that did not make it into most of the coverage. Describing its own cybersecurity results, the Beijing…
Category: AI
Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research
arXiv:2608.15052v1 Announce Type: new Abstract: Andy is an autonomous mathematical research agent that solves and verifies submitted problems, formulates…
T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework
arXiv:2608.14953v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code…
Demand-Driven Vertiport Siting and Discrete-Event Fleet Simulation for On-Demand Urban Air Mobility Network Design
arXiv:2608.14974v1 Announce Type: new Abstract: This paper presents a demand-driven framework for on-demand Urban Air Mobility (UAM) network design that…
RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder
arXiv:2608.14947v1 Announce Type: new Abstract: Detecting opioid craving from wearable physiological signals is critical yet difficult, with the potential…
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5
arXiv:2608.14992v1 Announce Type: new Abstract: Language-model systems increasingly read from stores they also write to, so a claim that was merely…
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
arXiv:2608.14945v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged…
Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid
arXiv:2608.14943v1 Announce Type: new Abstract: Agent skills are often injected in full on every request, increasing token cost. We compare four…
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks
arXiv:2608.14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but…
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
arXiv:2608.14940v1 Announce Type: new Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count…
