arXiv:2608.11233v1 Announce Type: cross Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent…
Category: AI
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
arXiv:2608.11232v1 Announce Type: cross Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data…
The Architecture Test: How to Tell Real Agentic AI From Rebadged Automation
Open pretty much any marketing technology vendor’s homepage today, and you will find the same three words somewhere above the fold: “powered by AI.” It…
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
arXiv:2608.11236v1 Announce Type: cross Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements…
Hospitals Adopted AI Before They Understood What They Were Adopting
Hospitals did not adopt artificial intelligence in a single, deliberate moment. It arrived in pieces: an imaging algorithm, a documentation assistant, a…
Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models
arXiv:2608.12304v1 Announce Type: new Abstract: Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking…
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
arXiv:2608.12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet…
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
arXiv:2608.12249v1 Announce Type: new Abstract: Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the…
Top AI lab researchers warned about automated AI research, and several of their predicted milestones have already fallen
IAPS fellow Severin Field interviewed 25 researchers from OpenAI, Anthropic, Google Deepmind, Meta, and US universities about recursive self-improvement.…
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings
arXiv:2608.12133v1 Announce Type: new Abstract: Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables,…
