arXiv:2608.11236v1 Announce Type: cross Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements…
Category: cs.AI updates on arXiv.org
Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models
arXiv:2608.12304v1 Announce Type: new Abstract: Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking…
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
arXiv:2608.12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet…
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
arXiv:2608.12249v1 Announce Type: new Abstract: Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the…
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings
arXiv:2608.12133v1 Announce Type: new Abstract: Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables,…
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
arXiv:2608.12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets. External oracles…
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We…
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
arXiv:2608.11994v1 Announce Type: new Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through…
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
arXiv:2608.12097v1 Announce Type: new Abstract: Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to…
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their…
