12 posts published in the last hour 05:32From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling 05:32Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs 05:32Breaking the 1.58-bit Barrier for Ternary LLMs 05:32Skill-based Agentic Evaluation for Real-time Data Science…
From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling
arXiv:2609.16493v1 Announce Type: new Abstract: Traditional catastrophe (CAT) risk models rely on costly manual construction to generate extreme weather…
Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs
arXiv:2609.16454v1 Announce Type: new Abstract: Recent work by Doshi and Hauser (2024), Bisbee et al. (2024), and Xie et al. (2026) raises concerns that…
Breaking the 1.58-bit Barrier for Ternary LLMs
arXiv:2609.16338v1 Announce Type: new Abstract: Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost…
Skill-based Agentic Evaluation for Real-time Data Science Tasks
arXiv:2609.16487v1 Announce Type: new Abstract: We present a framework for evaluating data-science agents on live, continuously updated data using…
Cross-Anatomy Transfer Versus Sparse Interpolation in Digital-Twin-Oriented Aortic Fluid-Structure Interaction Surrogates
arXiv:2609.16322v1 Announce Type: new Abstract: Surrogate credibility for fluid-structure interac- tion (FSI) requires distinguishing transfer across…
The AI-Enabled Scientific Frontier
arXiv:2609.16258v1 Announce Type: new Abstract: As artificial intelligence’s capabilities improve, it is increasingly viewed as a general scientific…
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
arXiv:2609.16305v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use,…
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
arXiv:2609.16251v1 Announce Type: new Abstract: Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks…
CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine
arXiv:2609.16301v1 Announce Type: new Abstract: Medical knowledge evolves continuously, whereas the parametric knowledge encoded in large language models…
Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models
Nums AI has released Causilo, a pretrained tabular foundation model for classification and regression with a scikit-learn interface. It posts the top…
Closing the Loop: Branch-and-Bound for Scalable Verification of Nonlinear Neural Feedback Systems
arXiv:2609.16298v1 Announce Type: new Abstract: Despite recent advances in the verification of nonlinear neural feedback systems, scalability remains the…
AI News Brief Hourly Summary 2026-09-16 07h : 11 posts
11 posts published in the last hour 04:32Artificial intelligence and biosecurity: capabilities, threat pathways, and defense-in-depth governance 04:32Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI 04:32Metacognitive Steering: Learning the Structure of Scientific Judgment…
Artificial intelligence and biosecurity: capabilities, threat pathways, and defense-in-depth governance
arXiv:2609.16213v1 Announce Type: new Abstract: Artificial intelligence is reshaping biological research across an increasingly connected…
Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI
arXiv:2609.16232v1 Announce Type: new Abstract: Geospatial artificial intelligence (GeoAI) powered by large language models (LLMs) is expanding the…
Metacognitive Steering: Learning the Structure of Scientific Judgment
arXiv:2609.16245v1 Announce Type: new Abstract: Long-horizon scientific discovery requires agents to alternate between exploration, disciplined execution,…
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
arXiv:2609.16247v1 Announce Type: new Abstract: Large language models sometimes behave in ways resembling human emotional responses, and recent work has…
Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
arXiv:2609.16215v1 Announce Type: new Abstract: GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops,…
