arXiv:2608.20735v1 Announce Type: new Abstract: Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action…
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design
arXiv:2608.20755v1 Announce Type: new Abstract: Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity…
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol
arXiv:2608.20729v1 Announce Type: new Abstract: Language-model agents can improve after failure or carry text across episodes without revising what counts…
VortexChat: An agentic framework for autonomous multi-objective integrated photonic design
arXiv:2608.20688v1 Announce Type: new Abstract: The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that…
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
arXiv:2608.20664v1 Announce Type: new Abstract: DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software…
Why2Speak: Faithful Reasoning for Abstaining Action Policies
arXiv:2608.20670v1 Announce Type: new Abstract: Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning…
CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery
arXiv:2608.20686v1 Announce Type: new Abstract: Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain…
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning
arXiv:2608.20717v1 Announce Type: new Abstract: Reliable confidence estimation is essential for using large language models in mathematical reasoning, but…
AI News Brief Hourly Summary 2026-08-24 09h : 11 posts
11 posts published in the last hour 06:32SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL 06:32Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work 06:32Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents…
SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL
arXiv:2608.20630v1 Announce Type: new Abstract: SQL systems increasingly expose AI functions for tasks such as classification, extraction, filtering,…
Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work
arXiv:2608.20622v1 Announce Type: new Abstract: Frontier models have collapsed the cost of writing custom code: a niche problem a specialist sees in their…
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents
arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring…
Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions
arXiv:2608.20649v1 Announce Type: new Abstract: Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces,…
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
arXiv:2608.20661v1 Announce Type: new Abstract: Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in…
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
arXiv:2608.20611v1 Announce Type: new Abstract: Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over…
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under…
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
arXiv:2608.20614v1 Announce Type: new Abstract: Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and…
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
arXiv:2608.20574v1 Announce Type: new Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or…
