arXiv:2606.12040v3 Announce Type: replace Abstract: The design of reinforced concrete (RC) highway barriers is a safety-critical engineering task that…
Category: AI
Rescaling Confidence: What Scale Design Reveals About LLM Metacognition
arXiv:2603.09309v3 Announce Type: replace Abstract: Verbalized confidence, in which LLMs report a numerical certainty score, is widely used to estimate…
VeriSim: A Configurable Framework for Stress-Testing Medical AI Under Patient Communication Noise
arXiv:2604.10441v2 Announce Type: replace Abstract: Medical large language models are typically evaluated on idealized patient cases that do not reflect…
TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas
arXiv:2603.16448v3 Announce Type: replace Abstract: Text-to-SQL parsing has achieved remarkable progress under the Full Schema Assumption. However, this…
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
arXiv:2604.06015v2 Announce Type: replace Abstract: Instruction tuning is commonly assumed to endow language models with a domain-general ability to…
5 Python Techniques for Efficient Resource Orchestration
This article explains 5 Python techniques for efficient resource orchestration and sticks to what’s stable today, 3.11 and later for the core techniques,…
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
arXiv:2603.15976v2 Announce Type: replace Abstract: While LLMs have accelerated scientific code generation, comprehensively evaluating generated code…
Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage
arXiv:2504.20007v4 Announce Type: replace Abstract: This paper proposes a novel interdisciplinary framework for analyzing police body-worn camera (BWC)…
GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay
arXiv:2609.11923v1 Announce Type: cross Abstract: Counterfactual regret minimization (CFR) is one of the few large numerical workloads that still runs…
Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning
arXiv:2506.14045v2 Announce Type: replace Abstract: Developing agents capable of exploring, planning and learning in complex open-ended environments is a…
