Abstract: Recent advances in AI for scientific discovery enable molecular understandingand design, yet reasoning over incomplete chemical representations remainsunclear.Markush structures, which encode molecular families through variable R-groupplaceholders (\textit{R\textsubscript{1}}, \textit{R\textsubscript{2}}, \textit{X}, etc.), are ubiquitous in pharmaceutical patents and requiregrounding across molecular, textual, and chemical information.However, existing molecule-language benchmarks focus on fully specifiedmolecules, leaving R-group grounding largely unevaluated.We introduce R-GroundBench:, a diagnostic benchmark built from real patent Markushstructures, featuring a Multiple-Choice (VQA) track with controlled difficultyand modality splits, and an open-ended Generation track.Our results reveal a substantial gap between recognition andmolecular grounding.While models achieve over 90\% accuracy on Easy VQA, performance drops to56–66\% on Hard VQA when shortcuts are controlled.Chemical-domain VLMs also remain unreliable, achieving only 25.7–46.2\% on HardVQA despite domain-specific pretraining.Moreover, Generation Exact Match remains below 20\% for most models and below8\% when visual input is required.These findings reveal that current AI systems lack reliable grounding andexecution for Markush editing, highlighting challenges for AI-drivenscientific discovery.
Read the original article:
