SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

arXiv:2609.06192v1 Announce Type: new
Abstract: Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final
outputs alone does not establish whether their conclusions are scientifically supported. We formulate
evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims
supported by results and visualizations from the same run. We introduce SciRIGOR, an evaluation framework and
benchmark comprising 100 cases from scientific articles across six domains and 17 subfields. The framework
reconstructs typed evidence graphs, separates artifact fidelity from relational validity, and scores complete
claim-support paths while localizing the earliest unsupported relation. Source-grounded alternative paths
accommodate scientifically equivalent analyses and visualizations. We evaluate 11 agent/model configurations.
On full-benchmark runs, claims agree with faithful and unfaithful results at nearly identical rates (91.8%
versus 91.0%). Yet no system exceeds 62.6% on the soft evidence-chain score or 18.0% strict whole-chain
success. These findings show that internal coherence does not establish scientific correctness: evaluation
must verify support along the complete data-to-claim path.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: