Abstract: Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or conflicting evidence. We introduce MEDLEY-BENCH, an open benchmark comparing structured private self-review and analyst-conditioned social revision from a common solo baseline. We evaluated 35 models from 12 families on 130 instances using the Medley Metacognition Score (MMS) and four taxonomy-aligned, rubric-derived composites. MMS point estimates were not consistently ordered in the available within-family size or generation comparisons. Under the prespecified ipsative procedure, the Evaluation-mapped composite had the lowest relative rubric score in 30 of 35 models; Self-regulation was lowest in four models and Control in one. This rubric- and centering-dependent pattern may reflect model behavior, judge severity, data-pipeline effects, or their combination; it is not evidence of an absolute Evaluation deficit. In an exploratory adversarial analysis of 11 purposively selected models, sensitivity to manipulated consensus labels ranged from near zero to larger response shifts. A preliminary human rubric-application study of 24 paired vignettes found a mean composite difference of 0.727 (95% CI: 0.500-0.942). Across 24 cross-rated response items from 12 shared vignettes, quadratic-weighted inter-reviewer agreement was kappa = 0.389, and response-profile human-LLM convergence was rho = 0.637. This preprint reports MEDLEY-BENCH v1.0, the audited proof-of-concept release. A separately versioned v1.5 will rerun the full protocol with corrected social-summary rendering and stronger scoring reproducibility, experimental control, and uncertainty analysis. MEDLEY-BENCH complements accuracy-based evaluation by characterizing prompted belief revision under ambiguity and social disagreement.
Read the original article: