arXiv:2608.14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model…
Author: script
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
arXiv:2608.14669v1 Announce Type: new Abstract: Artificial intelligence systems applied to mathematics verify correctness but not novelty: an…
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
arXiv:2608.14694v1 Announce Type: new Abstract: Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless…
AI News Brief Hourly Summary 2026-08-18 09h : 10 posts
10 posts published in the last hour 06:32When Uncertainty Isn’t Enough: An Empirical Study of Self-Correction in Code Generation 06:32Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring 06:32Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers…
When Uncertainty Isn’t Enough: An Empirical Study of Self-Correction in Code Generation
arXiv:2608.14659v1 Announce Type: new Abstract: Large language models for code generation often produce incorrect solutions without reliable indicators of…
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring
arXiv:2608.14666v1 Announce Type: new Abstract: Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that…
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
arXiv:2608.14641v1 Announce Type: new Abstract: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually…
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
arXiv:2608.14651v1 Announce Type: new Abstract: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency…
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
arXiv:2608.14667v1 Announce Type: new Abstract: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet…
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
arXiv:2608.14631v1 Announce Type: new Abstract: As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large…
