Evaluation Rigor from Graph Neural Networks to Graph Foundation Models: A Systematic Review and a Four-Axis Reporting Standard

Sergei Olegovich Kurashkin, В С Тынченко, А. С. Бородулин, Ahmad Hammoud, Tee Connie · Machine Learning and Knowledge Extraction · 2026

Graph machine learning reports steady progress across node, graph, and link prediction, across temporal and hypergraph frontiers, and across the emerging class of graph foundation models. This review asks a prior question: when a method is reported to outperform the alternatives, how far does the evidence support the claim? We organize the answer around four axes of evaluation rigor: statistical rigor (seeds, dispersion, formal significance testing), baseline fairness (budget-parity tuning of trivial and structure-agnostic baselines), data integrity (leakage, duplication, negative sampling, contamination), and claim integrity (whether gains survive fair tuning and discriminative benchmarks). Drawing on a criterion-based corpus of 150 studies, of which 51 were read in full depth, we find a consistent picture. Only two of 25 methodologically central backbone studies apply a formal between-method significance test, and reported gains repeatedly shrink or disappear once a trivial baseline is tuned to parity, a leaked split is repaired, or a pretrained model is evaluated on unseen data. We argue that these failures share one cause: the saturation of benchmarks that can no longer discriminate between methods. The principal output is a minimum reporting standard, a concrete four-axis checklist that authors and reviewers can apply at negligible cost.

Read the paper · More papers on PaperTik