arXiv:2608.20688v1 Announce Type: new Abstract: The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that…
Author: script
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
arXiv:2608.20664v1 Announce Type: new Abstract: DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software…
Why2Speak: Faithful Reasoning for Abstaining Action Policies
arXiv:2608.20670v1 Announce Type: new Abstract: Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning…
CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery
arXiv:2608.20686v1 Announce Type: new Abstract: Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain…
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning
arXiv:2608.20717v1 Announce Type: new Abstract: Reliable confidence estimation is essential for using large language models in mathematical reasoning, but…
AI News Brief Hourly Summary 2026-08-24 09h : 11 posts
11 posts published in the last hour 06:32SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL 06:32Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work 06:32Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents…
SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL
arXiv:2608.20630v1 Announce Type: new Abstract: SQL systems increasingly expose AI functions for tasks such as classification, extraction, filtering,…
Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work
arXiv:2608.20622v1 Announce Type: new Abstract: Frontier models have collapsed the cost of writing custom code: a niche problem a specialist sees in their…
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents
arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring…
Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions
arXiv:2608.20649v1 Announce Type: new Abstract: Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces,…
