arXiv:2608.26696v1 Announce Type: new Abstract: Enterprise deployments of autonomous AI agents inherit a control model built for human users and…
Tag: cs.AI updates on arXiv.org
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
arXiv:2608.26546v1 Announce Type: new Abstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world…
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured,…
ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving
arXiv:2608.26334v1 Announce Type: new Abstract: Automated theorem proving offers a natural foundation for recursive self-improvement in scientific…
Don’t Overthink, Don’t Underthink: Toward Adaptive Reasoning in Agentic AI
arXiv:2608.26442v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can…
Fine-Tuning of Transformer models with Frames
arXiv:2608.26430v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective…
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
arXiv:2608.26530v1 Announce Type: new Abstract: Long-horizon agent runs generate experience that can improve both the current run and future work. Most…
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
arXiv:2608.26535v1 Announce Type: new Abstract: Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning,…
Assessing mentalization in humans and large language models
arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization – the ability to infer others’ beliefs and intentions to guide one’s own choices – is a key…
SKILL.state: Scalable Long-Horizon Agent Skills
arXiv:2608.26263v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running…
