arXiv:2608.26696v1 Announce Type: new Abstract: Enterprise deployments of autonomous AI agents inherit a control model built for human users and…
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
arXiv:2608.26546v1 Announce Type: new Abstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world…
Always-on and self-starting AI agents might be OpenAI’s next big play
OpenAI is building a “Persistent Mode” for its AI agent Codex that stays active indefinitely and generates its own follow-up tasks. WIRED found the…
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured,…
ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving
arXiv:2608.26334v1 Announce Type: new Abstract: Automated theorem proving offers a natural foundation for recursive self-improvement in scientific…
Don’t Overthink, Don’t Underthink: Toward Adaptive Reasoning in Agentic AI
arXiv:2608.26442v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can…
Fine-Tuning of Transformer models with Frames
arXiv:2608.26430v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective…
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
arXiv:2608.26530v1 Announce Type: new Abstract: Long-horizon agent runs generate experience that can improve both the current run and future work. Most…
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context…
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
arXiv:2608.26535v1 Announce Type: new Abstract: Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning,…
AI News Brief Hourly Summary 2026-08-28 10h : 12 posts
12 posts published in the last hour 07:33Assessing mentalization in humans and large language models 07:33SKILL.state: Scalable Long-Horizon Agent Skills 07:336.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation 07:33FaithSieve:…
Assessing mentalization in humans and large language models
arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization – the ability to infer others’ beliefs and intentions to guide one’s own choices – is a key…
SKILL.state: Scalable Long-Horizon Agent Skills
arXiv:2608.26263v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running…
6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation
arXiv:2608.26236v1 Announce Type: new Abstract: We present a six-stage framework for auditing the reproducibility of scientific claims across a research…
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
arXiv:2608.26310v1 Announce Type: new Abstract: Large language models can now generate complex, multi-step mathematical proofs, but reliably determining…
Anthropic Opens 10,000 Free and Discounted Claude Seats for Scientists
Anthropic on Thursday, August 27, 2026, opened 10,000 free and discounted Claude Team subscription seats for scientists, announcing an expansion of its…
Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems
arXiv:2608.26306v1 Announce Type: new Abstract: A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is…
Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation – Identity Adequacy and Evidence Adequacy
arXiv:2608.26225v1 Announce Type: new Abstract: Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes,…
