arXiv:2608.26236v1 Announce Type: new Abstract: We present a six-stage framework for auditing the reproducibility of scientific claims across a research…
Tag: AI
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
arXiv:2608.26310v1 Announce Type: new Abstract: Large language models can now generate complex, multi-step mathematical proofs, but reliably determining…
Anthropic Opens 10,000 Free and Discounted Claude Seats for Scientists
Anthropic on Thursday, August 27, 2026, opened 10,000 free and discounted Claude Team subscription seats for scientists, announcing an expansion of its…
Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems
arXiv:2608.26306v1 Announce Type: new Abstract: A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is…
Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation – Identity Adequacy and Evidence Adequacy
arXiv:2608.26225v1 Announce Type: new Abstract: Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes,…
GameWAM: A World Action Model for Video Games
arXiv:2608.26200v1 Announce Type: new Abstract: Modern video games combine first-person perception, rapid visual changes, persistent world state, and…
The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts
arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment…
LLM Agents for Time-Series: A Survey
arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary…
Same Model, Different Harness: Different Coding-Agent Results
arXiv:2608.26218v1 Announce Type: new Abstract: A coding agent combines a model with a harness, which decides what the model sees, which tools it can use,…
Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs
arXiv:2608.26191v1 Announce Type: new Abstract: Incident risk prediction from longitudinal electronic health records (EHRs) is challenging because…
