arXiv:2608.16147v1 Announce Type: new Abstract: Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting…
Category: cs.AI updates on arXiv.org
Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
arXiv:2608.16094v1 Announce Type: new Abstract: Accurate protein structure prediction is fundamental to structural biology because protein structure…
TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes,…
ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction
arXiv:2608.15979v1 Announce Type: new Abstract: Large language models produce outputs presented as discoveries – new proofs, conjectures, or molecules.…
Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance
arXiv:2608.16055v1 Announce Type: new Abstract: Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within…
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same…
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
arXiv:2608.16084v1 Announce Type: new Abstract: Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic…
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment
arXiv:2608.15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal…
Solvable Sokoban Without a Solver via Diffusion
arXiv:2608.15958v1 Announce Type: new Abstract: Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be…
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
arXiv:2608.15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased…
