arXiv:2608.15101v1 Announce Type: new Abstract: Policy evaluation often estimates direct benefits and costs while treating the institutional environment…
Author: script
AI News Brief Hourly Summary 2026-08-18 13h : 14 posts
14 posts published in the last hour 10:33Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents 10:33GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG 10:33AI’s recursive self-improvement might not come so quickly after all 10:33TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning…
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
arXiv:2608.15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM)…
GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
arXiv:2608.15056v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or…
AI’s recursive self-improvement might not come so quickly after all
The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code,…
TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning
arXiv:2608.15055v1 Announce Type: new Abstract: Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while…
Claude Code gets a /design command that lets developers create UI mockups right in the terminal
With the /design command, Anthropic brings a visual design workflow directly into Claude Code. Developers can generate UI mockups as artboards right in…
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
arXiv:2608.15065v1 Announce Type: new Abstract: Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same…
We still don’t know how people are really using AI
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data…
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents
arXiv:2608.15064v1 Announce Type: new Abstract: Parsing visual documents into machine-readable representations is fundamental to document intelligence.…
