Future-proofing online exams amid rising demand and tightening privacy laws. The global online exam software market, valued at $9.4b in 2025, is expected…
Author: script
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models
arXiv:2608.15145v1 Announce Type: new Abstract: Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain…
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
arXiv:2608.15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may…
Validation-Frontier Representation Selection under Constrained Observation
arXiv:2608.15095v1 Announce Type: new Abstract: AI systems deployed outside clean benchmark settings often rely on observations that are incomplete,…
Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents
arXiv:2608.15109v1 Announce Type: new Abstract: Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical…
Anthropic’s per-token cost runs 4.4 times the average on Vercel, and developers keep paying
Anthropic dominated Vercel’s AI Gateway spending in July, pulling in 65.1 percent of total revenue while accounting for only 30 percent of tokens…
Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems
arXiv:2608.15082v1 Announce Type: new Abstract: Cold chain logistics has advanced technologically, yet most deployed systems remain reactive monitors, not…
Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial…
Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation
arXiv:2608.15101v1 Announce Type: new Abstract: Policy evaluation often estimates direct benefits and costs while treating the institutional environment…
AI News Brief Hourly Summary 2026-08-18 13h : 14 posts
14 posts published in the last hour 10:33Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents 10:33GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG 10:33AI’s recursive self-improvement might not come so quickly after all 10:33TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning…
