OpenAI is building a “Persistent Mode” for its AI agent Codex that stays active indefinitely and generates its own follow-up tasks. WIRED found the…
Tag: AI
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured,…
ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving
arXiv:2608.26334v1 Announce Type: new Abstract: Automated theorem proving offers a natural foundation for recursive self-improvement in scientific…
Don’t Overthink, Don’t Underthink: Toward Adaptive Reasoning in Agentic AI
arXiv:2608.26442v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can…
Fine-Tuning of Transformer models with Frames
arXiv:2608.26430v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective…
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
arXiv:2608.26530v1 Announce Type: new Abstract: Long-horizon agent runs generate experience that can improve both the current run and future work. Most…
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context…
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
arXiv:2608.26535v1 Announce Type: new Abstract: Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning,…
Assessing mentalization in humans and large language models
arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization – the ability to infer others’ beliefs and intentions to guide one’s own choices – is a key…
SKILL.state: Scalable Long-Horizon Agent Skills
arXiv:2608.26263v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running…
