Piloting the world’s first double-blind AI evaluations
Author: script
SWE-Prime: Fewer Trajectories, Better Performance
arXiv:2608.27449v1 Announce Type: cross Abstract: To improve large language models’ ability to resolve real-world software issues, prior work has focused…
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
arXiv:2608.27439v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can…
From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
arXiv:2608.27442v1 Announce Type: cross Abstract: In real-world software development, code review typically involves iterative interactions between…
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit
arXiv:2608.27427v1 Announce Type: cross Abstract: Large language model (LLM) agents in governed organizations must let the persona (instructions, tone,…
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
arXiv:2608.27406v1 Announce Type: cross Abstract: State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment,…
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
arXiv:2608.27424v1 Announce Type: cross Abstract: Static scanners are increasingly used to identify executable or otherwise unsafe content in machine-…
Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions
arXiv:2608.27392v1 Announce Type: cross Abstract: Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but…
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
arXiv:2608.27395v1 Announce Type: cross Abstract: Video carries the temporal structure of the physical world, yet learning representations from it has…
Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models
arXiv:2608.27367v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision…
