arXiv:2608.12304v1 Announce Type: new Abstract: Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking…
Author: script
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
arXiv:2608.12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet…
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
arXiv:2608.12249v1 Announce Type: new Abstract: Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the…
Top AI lab researchers warned about automated AI research, and several of their predicted milestones have already fallen
IAPS fellow Severin Field interviewed 25 researchers from OpenAI, Anthropic, Google Deepmind, Meta, and US universities about recursive self-improvement.…
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings
arXiv:2608.12133v1 Announce Type: new Abstract: Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables,…
IBM Embeds OpenAI Models Into Its Consulting Delivery Platform
IBM announced a strategic partnership with OpenAI on August 13, 2026, that embeds OpenAI’s frontier models, including GPT-5.6, and products such as Codex…
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
arXiv:2608.12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets. External oracles…
Fable 5’s slow adoption suggests corporate willingness to pay for frontier AI has hit a ceiling
Anthropic’s Fable 5 is considered the most powerful AI model on the market, but U.S. companies are barely buying it. According to Ramp data, Fable 5…
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We…
AI News Brief Hourly Summary 2026-08-13 13h : 11 posts
11 posts were published in the last hour 10:32 : Claim-Level Reliability Assessment for Efficient Test-Time Reasoning 10:32 : Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges 10:32 : Mechanist: AI as a Scientific Instrument for Discovering…
