15 posts published in the last hour 19:32Assessing LLMs’ mathematical abilities requires understanding the various mechanisms of mathematical creativity 19:32FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection 19:32When Single-Dataset Conclusions Fail: A 45-Task Study…
Assessing LLMs’ mathematical abilities requires understanding the various mechanisms of mathematical creativity
arXiv:2608.16118v1 Announce Type: new Abstract: How should we assess whether large language models can perform mathematical invention? I argue that this…
FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
arXiv:2608.16148v1 Announce Type: new Abstract: Multi-view multi-label feature selection aims to identify a compact and informative feature subset from…
When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification
arXiv:2608.16147v1 Announce Type: new Abstract: Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting…
Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
arXiv:2608.16094v1 Announce Type: new Abstract: Accurate protein structure prediction is fundamental to structural biology because protein structure…
Gemini in Chrome Opens to All U.S. Android Users as Auto Browse Goes Mobile
Google opened Gemini in Chrome to all Android users in the United States on August 18, 2026, bringing its built-in browsing assistant to the full U.S.…
TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes,…
ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction
arXiv:2608.15979v1 Announce Type: new Abstract: Large language models produce outputs presented as discoveries – new proofs, conjectures, or molecules.…
Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance
arXiv:2608.16055v1 Announce Type: new Abstract: Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within…
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same…
Amazon Bedrock AgentCore payments is now generally available: Enabling agents to transact safely and autonomously at scale
Amazon Bedrock AgentCore payments is now generally available, enabling AI agents to autonomously transact at scale with built-in spending guardrails,…
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
arXiv:2608.16084v1 Announce Type: new Abstract: Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic…
OpenAI says it’s “pacing model development” as AI cybersecurity risks grow too dangerous
OpenAI is deliberately “pacing AI model development,” partly because the upcoming “Astra” model may be close to gaining critical cyberattack capabilities.…
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment
arXiv:2608.15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal…
AI News Brief Hourly Summary 2026-08-18 21h : 17 posts
17 posts published in the last hour 18:32Solvable Sokoban Without a Solver via Diffusion 18:32UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations 18:32Augmenting Text to Increase Translation Difficulty 18:32How Much Memory Does Your Agent Actually Need? 18:32Navigation-Informed Embeddings: Dense-Retriever…
Solvable Sokoban Without a Solver via Diffusion
arXiv:2608.15958v1 Announce Type: new Abstract: Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be…
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
arXiv:2608.15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased…
