Piloting the world’s first double-blind AI evaluations
SWE-Prime: Fewer Trajectories, Better Performance
arXiv:2608.27449v1 Announce Type: cross Abstract: To improve large language models’ ability to resolve real-world software issues, prior work has focused…
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
arXiv:2608.27439v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can…
From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
arXiv:2608.27442v1 Announce Type: cross Abstract: In real-world software development, code review typically involves iterative interactions between…
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit
arXiv:2608.27427v1 Announce Type: cross Abstract: Large language model (LLM) agents in governed organizations must let the persona (instructions, tone,…
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
arXiv:2608.27406v1 Announce Type: cross Abstract: State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment,…
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
arXiv:2608.27424v1 Announce Type: cross Abstract: Static scanners are increasingly used to identify executable or otherwise unsafe content in machine-…
Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions
arXiv:2608.27392v1 Announce Type: cross Abstract: Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but…
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
arXiv:2608.27395v1 Announce Type: cross Abstract: Video carries the temporal structure of the physical world, yet learning representations from it has…
Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models
arXiv:2608.27367v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision…
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
arXiv:2608.27397v1 Announce Type: cross Abstract: Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts…
OpenAI to start showing ads on ChatGPT’s free and Go tiers in India
OpenAI has more than 100 million weekly active ChatGPT users in India, a huge chunk of whom are on the free or the lower-priced Go tiers.
How Language Models Organize and Structure Moral Knowledge
arXiv:2608.27402v1 Announce Type: cross Abstract: How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but…
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
arXiv:2608.27345v1 Announce Type: cross Abstract: Recent video generation models are increasingly framed as world models. Many physical processes can…
Stageboost: Recommending Signals Based on Counterfactual Estimation
arXiv:2608.27366v1 Announce Type: cross Abstract: Signals are short textual or visual snippets displayed on the eBay View-Item (VI) page, providing…
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
arXiv:2608.27309v1 Announce Type: cross Abstract: Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs…
KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations
arXiv:2608.27365v1 Announce Type: cross Abstract: Physics-integrated 3D Gaussian representations now allow reconstructed deformable objects to be…
RCMN: Understanding Misleadingness in Influential Public Discourse
arXiv:2608.27358v1 Announce Type: cross Abstract: Influential public discourse shapes public beliefs and can also mislead, not only through what is…
