arXiv:2605.29486v2 Announce Type: replace-cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments…
Category: AI
SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models
arXiv:2605.25420v2 Announce Type: replace-cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource…
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
arXiv:2604.09508v2 Announce Type: replace-cross Abstract: Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and…
An InSAR Phase Unwrapping Framework for Large-scale and Complex Events
arXiv:2603.21378v2 Announce Type: replace-cross Abstract: Phase unwrapping remains a critical and challenging problem in InSAR processing, particularly in…
Early Stopping for Large Reasoning Models via Confidence Dynamics
arXiv:2604.04930v2 Announce Type: replace-cross Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but…
A Systematic Comparison of Training Objectives for Out-of-Distribution Detection in Image Classification
arXiv:2603.07571v3 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is critical in safety-sensitive applications. While this…
Adaptive Stopping for Multi-Turn LLM Reasoning
arXiv:2604.01413v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as…
Nous Research Ships Bot Mode for Hermes Agent, Turning Agent Profiles Into a Roster of Named Bots
Nous Research has shipped Bot Mode for Hermes Agent, its MIT-licensed open source agent. Bot Mode replaces the single-agent session list with a roster of…
ArGEnT: Arbitrary Geometry-encoded Transformer for Operator Learning
arXiv:2602.11626v3 Announce Type: replace-cross Abstract: Learning solution operators on arbitrary geometries remains a central challenge in scientific…
A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models
arXiv:2602.09992v2 Announce Type: replace-cross Abstract: Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using…
