arXiv:2608.07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count—a…
Tag: cs.AI updates on arXiv.org
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model…
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
arXiv:2608.07367v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly used as a primary source of information and advice,…
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
arXiv:2608.07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous…
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
arXiv:2608.07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency:…
CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
arXiv:2608.07424v1 Announce Type: new Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions,…
FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
arXiv:2608.07400v1 Announce Type: new Abstract: Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible…
QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting
arXiv:2608.07363v1 Announce Type: new Abstract: Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility…
Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons
arXiv:2608.07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets — tens of seconds rather than hours — are…
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
arXiv:2608.07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models…
