arXiv:2609.05258v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to formulate optimization models from…
Category: cs.AI updates on arXiv.org
Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets
arXiv:2609.05194v1 Announce Type: cross Abstract: The number of discrete class-separability jumps observed during ResNet finetuning is examined…
A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR
arXiv:2609.05221v1 Announce Type: cross Abstract: Large language models (LLMs) show strong reasoning ability, but their explanations can remain…
CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls
arXiv:2609.05269v1 Announce Type: cross Abstract: LLM agent systems increasingly combine provenance tracking, authorization, policy enforcement, protocol…
A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
arXiv:2609.05143v1 Announce Type: cross Abstract: The integration of artificial intelligence (AI), particularly large language models (LLMs), into…
A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning
arXiv:2609.05133v1 Announce Type: cross Abstract: This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy…
AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics
arXiv:2609.05157v1 Announce Type: cross Abstract: Formalizing mathematics in a proof assistant, where a machine checks every definition, statement and…
Beyond Stationarity in Time Series: Discovering Causal Structures and Latent Regimes via Markov Blankets
arXiv:2609.05150v1 Announce Type: cross Abstract: This paper introduces Regime-aware Constraint-Based and Noise-Based causal discovery with Markov…
TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors
arXiv:2609.05117v1 Announce Type: cross Abstract: Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful…
Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection
arXiv:2609.05066v1 Announce Type: cross Abstract: As the scale of video surveillance data outpaces manual annotation capacities, weakly supervised video…
