arXiv:2608.07367v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly used as a primary source of information and advice,…
Category: cs.AI updates on arXiv.org
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
arXiv:2608.07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous…
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
arXiv:2608.07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency:…
CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
arXiv:2608.07424v1 Announce Type: new Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions,…
FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
arXiv:2608.07400v1 Announce Type: new Abstract: Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible…
QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting
arXiv:2608.07363v1 Announce Type: new Abstract: Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility…
Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons
arXiv:2608.07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets — tens of seconds rather than hours — are…
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
arXiv:2608.07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models…
Curriculum as Code: An AI-Assisted Architecture for Instructional Design in STEM Education
arXiv:2608.07364v1 Announce Type: new Abstract: Contribution: This paper presents a six-phase AI-assisted instructional design architecture based on the…
An End-to-End Agent Auditing Engine
arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure…