arXiv:2608.16118v1 Announce Type: new Abstract: How should we assess whether large language models can perform mathematical invention? I argue that this…
Author: script
FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
arXiv:2608.16148v1 Announce Type: new Abstract: Multi-view multi-label feature selection aims to identify a compact and informative feature subset from…
When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification
arXiv:2608.16147v1 Announce Type: new Abstract: Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting…
Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
arXiv:2608.16094v1 Announce Type: new Abstract: Accurate protein structure prediction is fundamental to structural biology because protein structure…
Gemini in Chrome Opens to All U.S. Android Users as Auto Browse Goes Mobile
Google opened Gemini in Chrome to all Android users in the United States on August 18, 2026, bringing its built-in browsing assistant to the full U.S.…
TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes,…
ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction
arXiv:2608.15979v1 Announce Type: new Abstract: Large language models produce outputs presented as discoveries – new proofs, conjectures, or molecules.…
Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance
arXiv:2608.16055v1 Announce Type: new Abstract: Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within…
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same…
