Comparative Evaluation of Feature Extractors, Aggregation Strategies, and Classification Hierarchies for Ovarian Cancer Subtype Classification in Whole Slide Images
Ho-Jung Song, You Sang Cho, Yong Suk Kim · Diagnostics · 2026
Background/Objectives: Multiple instance learning (MIL) is widely used for automated classification of epithelial ovarian cancer subtypes from whole slide images (WSIs), but the relative contributions of feature extractor, aggregation strategy, and classification framework (flat vs. hierarchical) choices remain unclear under severe class imbalance. Methods: We evaluated 36 configurations on 510 WSIs from the UBC-OCEAN dataset using stratified five-fold cross-validation, comparing three pathology foundation models (Phikon-v2, CTransPath, UNI), six aggregators (mean/max pooling, ABMIL, CLAM-SB, DSMIL, DTP-TransMIL), and two classification strategies. Pathologist-annotated WSIs assessed attention map interpretability. Results: Feature extractor selection contributed substantially more variance than aggregator choice. Cascade balanced accuracy ranged from 0.538 (Phikon-v2) to 0.925 (UNI); CTransPath (~32 K pretraining WSIs) reached 0.870, exceeding Phikon-v2 (~58 K WSIs) and approaching UNI (~100 K+ WSIs), indicating that pretraining objective and architecture contribute as substantially as scale. The hierarchical cascade consistently improved high-grade serous carcinoma (HGSC) recall across all six evaluated configurations (+0.073 to +0.530), detecting 206 of 217 cases (0.949) with UNI max pooling. Quantitative spatial alignment analysis confirmed that both stronger feature extractors—CTransPath and UNI—generated significantly more spatially structured attention distributions than Phikon-v2 (paired Wilcoxon, p = 0.008 and p = 0.032, respectively). Conclusions: Feature extractor choice contributed more variance than aggregator selection, with the largest gap between Phikon-v2 and stronger extractors. Hierarchical cascades consistently improved HGSC recall across all configurations.