arXiv:2608.17823v1 Announce Type: cross Abstract: Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet…
Category: cs.AI updates on arXiv.org
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
arXiv:2608.17719v1 Announce Type: cross Abstract: Context: Software systems that depend on commercial large language model APIs must migrate to successor…
Training with synthetic data for drone detection in thermal imagery
arXiv:2608.17799v1 Announce Type: cross Abstract: Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging…
MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps
arXiv:2608.17659v1 Announce Type: cross Abstract: LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research…
GADR: Gathering Architecture Decision Records from Meeting Transcriptions
arXiv:2608.17694v1 Announce Type: cross Abstract: Existing LLM-based approaches to Architecture Decision Record (ADR) generation share a critical and…
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models
arXiv:2608.17715v1 Announce Type: cross Abstract: Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to…
Benchmarking Automated Security Patch Backporting: How Far Are We?
arXiv:2608.17671v1 Announce Type: cross Abstract: Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools…
Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees
arXiv:2608.17703v1 Announce Type: cross Abstract: Mobile robots that operate in side by side with humans and critical facilities must reach their goals at…
From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support
arXiv:2608.17618v1 Announce Type: cross Abstract: Learning analytics models can identify students at risk of poor performance, but they do not directly…
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
arXiv:2608.17597v1 Announce Type: cross Abstract: Large language models are increasingly deployed through agent harnesses that manage tools, extensions,…
