arXiv:2609.00366v1 Announce Type: cross Abstract: High test accuracy and good aggregate calibration do not show whether an individual prediction is…
Tag: cs.AI updates on arXiv.org
Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
arXiv:2609.00330v1 Announce Type: cross Abstract: In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether…
Latent-Space No-Arbitrage Geometry of Generative Models for Implied Volatility Surfaces
arXiv:2609.00332v1 Announce Type: cross Abstract: Generative models for implied volatility surfaces must produce outputs that satisfy static no-arbitrage…
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
arXiv:2609.00351v1 Announce Type: cross Abstract: Large language models can hide hidden behaviors that activate only under narrow conditions, such as…
The Curse of Multilinguality in Lexical Normalization
arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u,…
A Human-AI Theorem Connecting Spontaneous and Field-Induced Mechanisms of Collective Behavior in One Dimension
arXiv:2609.00322v1 Announce Type: cross Abstract: Can an artificial intelligence (AI) generate a scientific hypothesis outside a human collaborator’s…
Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains
arXiv:2609.00297v1 Announce Type: cross Abstract: Solving multiphysics partial differential equations (PDEs) remains a major challenge in scientific…
WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network Optimization
arXiv:2609.00284v1 Announce Type: cross Abstract: Emerging 6G wireless networks are expected to operate across diverse deployment scenarios, where…
Workload Identification with Physical Side Channels for AI Governance
arXiv:2609.00309v1 Announce Type: cross Abstract: AI compute verification is one of the first tangible and tractable points for international policy aimed…
Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer’s Disease Detection
arXiv:2609.00276v1 Announce Type: cross Abstract: Speech-based Alzheimer’s disease (AD) detection increasingly relies on speech-enhanced and curated…
