arXiv:2608.20666v1 Announce Type: cross Abstract: While Digital sky surveys provide excellent throughput of image data and can cover a large footprint,…
Category: cs.AI updates on arXiv.org
C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination
arXiv:2608.20667v1 Announce Type: cross Abstract: Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and…
RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction
arXiv:2608.20656v1 Announce Type: cross Abstract: Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting…
ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection
arXiv:2608.20637v1 Announce Type: cross Abstract: Static analyzers have been widely adopted for vulnerability detection in C/C++ programs. Query-based…
Testing and Evaluation of Agentic AI Systems In Military Command and Control
arXiv:2608.20597v1 Announce Type: cross Abstract: Agentic AI systems are being procured for military command and control (C2) under public commitments to…
JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification
arXiv:2608.20607v1 Announce Type: cross Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings,…
When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation
arXiv:2608.20627v1 Announce Type: cross Abstract: Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation…
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
arXiv:2608.20634v1 Announce Type: cross Abstract: Agents learn to act through interaction with environments, yet the environments used for training are…
ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations
arXiv:2608.20539v1 Announce Type: cross Abstract: Digital twin simulations show promise, but current empirical evidence suggests that the approach should…
Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes
arXiv:2608.20521v1 Announce Type: cross Abstract: Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with…
