arXiv:2608.14771v1 Announce Type: new Abstract: Making language models solve constraint problems reliably often means having them translate the problem…
Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition
arXiv:2608.14673v1 Announce Type: new Abstract: Chapter 6 of OpenAI’s *Ten Advances in Mathematics and Theoretical Computer Science* claims an exponential…
Synchronized Logit Steering: Real-world Steganography
arXiv:2608.14697v1 Announce Type: new Abstract: Steganography in large language models offers a way to embed hidden messages within natural-sounding text.…
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
arXiv:2608.14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model…
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
arXiv:2608.14669v1 Announce Type: new Abstract: Artificial intelligence systems applied to mathematics verify correctness but not novelty: an…
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
arXiv:2608.14694v1 Announce Type: new Abstract: Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless…
AI News Brief Hourly Summary 2026-08-18 09h : 10 posts
10 posts published in the last hour 06:32When Uncertainty Isn’t Enough: An Empirical Study of Self-Correction in Code Generation 06:32Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring 06:32Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers…
When Uncertainty Isn’t Enough: An Empirical Study of Self-Correction in Code Generation
arXiv:2608.14659v1 Announce Type: new Abstract: Large language models for code generation often produce incorrect solutions without reliable indicators of…
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring
arXiv:2608.14666v1 Announce Type: new Abstract: Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that…
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
arXiv:2608.14641v1 Announce Type: new Abstract: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually…
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
arXiv:2608.14651v1 Announce Type: new Abstract: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency…
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
arXiv:2608.14667v1 Announce Type: new Abstract: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet…
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
arXiv:2608.14631v1 Announce Type: new Abstract: As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large…
Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP
arXiv:2608.14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic’s Model Context…
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting.…
Learning Agent Execution for KV-Cache Management in Agentic Serving
arXiv:2608.14624v1 Announce Type: new Abstract: Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user…
Large Language Models and their Awareness of Mechanics and Spatial Geometry
arXiv:2608.14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning…
AI News Brief Hourly Summary 2026-08-18 08h : 11 posts
11 posts published in the last hour 05:32When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning 05:32Position: Medical AI Neglects Real Treatment Outcomes 05:32The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM…
