arXiv:2609.26826v1 Announce Type: cross Abstract: Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not…
Tag: AI
G\”odel’s and Scott’s Variants of the Ontological Argument in Lean 4
arXiv:2609.26806v1 Announce Type: cross Abstract: This paper presents a complete, structure-preserving port to Lean 4 of the Isabelle/HOL dataset…
Learning Stiffness Dependent Fluid Structure Dynamics from Coarse Flow Representations
arXiv:2609.26816v1 Announce Type: cross Abstract: This paper develops a data-driven framework for long-term prediction of fluid–structure interaction…
Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management
arXiv:2609.26828v1 Announce Type: cross Abstract: NAND-backed storage offers the capacity needed to scale LLM prefix caching, but its block I/O path…
Signal2Symbol: Neuro-Symbolic Temporal Reasoning for Explainable Physiological Time-Series Anomaly Detection
arXiv:2609.26820v1 Announce Type: cross Abstract: Physiological time series such as electrocardiograms (ECG) and electroencephalograms (EEG) exhibit…
StudentBench: AI and human tutoring yield equivalent GRE learning gains
arXiv:2609.28470v1 Announce Type: new Abstract: Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at…
Shutdown Sabotage Propensities in Multi-Agent Systems
arXiv:2609.28274v1 Announce Type: new Abstract: The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been…
Attention-based representations for multi-task computation
arXiv:2608.04243v1 Announce Type: cross Abstract: Multi-head attention layers produce vector representations that support multiple downstream tasks. We…
An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act’s Code of Practice
arXiv:2609.28335v1 Announce Type: new Abstract: Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or…
Learning the Cost of Reliable Inference
arXiv:2609.28322v1 Announce Type: new Abstract: Benchmarking and routing platforms increasingly act as intermediaries connecting large language model…
