arXiv:2608.14765v1 Announce Type: new Abstract: Data cleaning without a trusted clean reference is challenging because unusual values may represent either…
Category: AI
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
arXiv:2608.14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current…
Advanced modelling and data analytics in aviation
arXiv:2608.14746v1 Announce Type: new Abstract: The aviation industry characterized by its stringent safety standards has seen a growing need for…
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
arXiv:2608.14707v1 Announce Type: new Abstract: As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents…
Haut.AI Expands Across 4,000 O Boticário Stores After Pilot Lifts Skincare Order Value 80%
Estonian skin-analysis vendor Haut.AI is moving from a 24-store experiment to a national retail deployment. On August 18, 2026, the company and Brazil’s…
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
arXiv:2608.14771v1 Announce Type: new Abstract: Making language models solve constraint problems reliably often means having them translate the problem…
Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition
arXiv:2608.14673v1 Announce Type: new Abstract: Chapter 6 of OpenAI’s *Ten Advances in Mathematics and Theoretical Computer Science* claims an exponential…
Synchronized Logit Steering: Real-world Steganography
arXiv:2608.14697v1 Announce Type: new Abstract: Steganography in large language models offers a way to embed hidden messages within natural-sounding text.…
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
arXiv:2608.14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model…
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
arXiv:2608.14669v1 Announce Type: new Abstract: Artificial intelligence systems applied to mathematics verify correctness but not novelty: an…
