arXiv:2608.20970v1 Announce Type: new Abstract: Recent work in mechanistic interpretability has studied how large language models recall facts stored in…
Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning
arXiv:2608.20960v1 Announce Type: new Abstract: Language models (LMs) are trained on static scientific corpora, whereas scientific knowledge continuously…
AI chatbots regularly link pregnant users to anti-abortion websites without disclosure
When asked about unplanned pregnancies, AI chatbots regularly link to anti-abortion groups without disclosing their stance. In an AlgorithmWatch…
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
arXiv:2608.20958v1 Announce Type: new Abstract: E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where…
AI News Brief Hourly Summary 2026-08-24 12h : 12 posts
12 posts published in the last hour 09:33No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators 09:33UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists 09:33Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control 09:33ReCurveflow: A Flow Matching…
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators
arXiv:2608.20938v1 Announce Type: new Abstract: Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems…
UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists
arXiv:2608.20918v1 Announce Type: new Abstract: Organizations maintain task-specific adapters for open-weight language models, and each new base-model…
Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control
arXiv:2608.20936v1 Announce Type: new Abstract: World models for continuous control are commonly trained for a fixed physical system and can degrade when…
ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries
arXiv:2608.20869v1 Announce Type: new Abstract: Predicting transition states (TS) in chemical reactions is crucial, as they provide insights into reaction…
The Logic of Machine Self-Preservation
arXiv:2608.20940v1 Announce Type: new Abstract: There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation,…
Foundation Models for Partial Causal Identification
arXiv:2608.20841v1 Announce Type: new Abstract: This paper investigates the development of causal foundation models for bounding the effect of…
MGAL: A Multilingual Granularity-Aware Long-Context Benchmark
arXiv:2608.20853v1 Announce Type: new Abstract: Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing…
TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding
arXiv:2608.20844v1 Announce Type: new Abstract: Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often…
RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation
arXiv:2608.20845v1 Announce Type: new Abstract: Nearly every retrieval-augmented question-answering system in production ships with a hidden interpreter:…
Nvidia in talks to invest in Perplexity at $30 billion-plus valuation
Nvidia is negotiating an investment in Perplexity at a valuation above $30 billion, more than 50 percent higher than its last funding round, The…
Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems
arXiv:2608.20864v1 Announce Type: new Abstract: Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its…
AI News Brief Hourly Summary 2026-08-24 11h : 11 posts
11 posts published in the last hour 08:32SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers 08:32Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress 08:32Dynamic Context Scheduling: Learning Beyond the Static…
SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
arXiv:2608.20802v1 Announce Type: new Abstract: Human motion forecasters are increasingly accurate and fast, but reliable deployment requires uncertainty…
