arXiv:2608.15095v1 Announce Type: new Abstract: AI systems deployed outside clean benchmark settings often rely on observations that are incomplete,…
Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents
arXiv:2608.15109v1 Announce Type: new Abstract: Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical…
Anthropic’s per-token cost runs 4.4 times the average on Vercel, and developers keep paying
Anthropic dominated Vercel’s AI Gateway spending in July, pulling in 65.1 percent of total revenue while accounting for only 30 percent of tokens…
Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems
arXiv:2608.15082v1 Announce Type: new Abstract: Cold chain logistics has advanced technologically, yet most deployed systems remain reactive monitors, not…
Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial…
Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation
arXiv:2608.15101v1 Announce Type: new Abstract: Policy evaluation often estimates direct benefits and costs while treating the institutional environment…
AI News Brief Hourly Summary 2026-08-18 13h : 14 posts
14 posts published in the last hour 10:33Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents 10:33GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG 10:33AI’s recursive self-improvement might not come so quickly after all 10:33TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning…
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
arXiv:2608.15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM)…
GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
arXiv:2608.15056v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or…
AI’s recursive self-improvement might not come so quickly after all
The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code,…
TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning
arXiv:2608.15055v1 Announce Type: new Abstract: Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while…
Claude Code gets a /design command that lets developers create UI mockups right in the terminal
With the /design command, Anthropic brings a visual design workflow directly into Claude Code. Developers can generate UI mockups as artboards right in…
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
arXiv:2608.15065v1 Announce Type: new Abstract: Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same…
We still don’t know how people are really using AI
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data…
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents
arXiv:2608.15064v1 Announce Type: new Abstract: Parsing visual documents into machine-readable representations is fundamental to document intelligence.…
S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
arXiv:2608.15018v1 Announce Type: new Abstract: Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory…
LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning
arXiv:2608.15041v1 Announce Type: new Abstract: Coordinating multiple interacting units in complex engineering systems is challenging when system…
SCOPE: Score-Isolated Agentic Optimization for Video World Models
arXiv:2608.15043v1 Announce Type: new Abstract: Video world models are increasingly used as simulators for planning and embodied decision making, yet…
