arXiv:2608.16055v1 Announce Type: new Abstract: Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within…
Category: AI
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same…
Amazon Bedrock AgentCore payments is now generally available: Enabling agents to transact safely and autonomously at scale
Amazon Bedrock AgentCore payments is now generally available, enabling AI agents to autonomously transact at scale with built-in spending guardrails,…
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
arXiv:2608.16084v1 Announce Type: new Abstract: Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic…
OpenAI says it’s “pacing model development” as AI cybersecurity risks grow too dangerous
OpenAI is deliberately “pacing AI model development,” partly because the upcoming “Astra” model may be close to gaining critical cyberattack capabilities.…
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment
arXiv:2608.15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal…
Solvable Sokoban Without a Solver via Diffusion
arXiv:2608.15958v1 Announce Type: new Abstract: Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be…
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
arXiv:2608.15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased…
Augmenting Text to Increase Translation Difficulty
arXiv:2608.15932v1 Announce Type: new Abstract: As state-of-the-art machine translation models saturate standard benchmarks, the field needs more…
