Coca-Mil: Attention-Based Handcrafted-Deep Feature Fusion in Computational Pathology
Paras Goel, Saarthak Kapse, Pushpak Pati, Prateek Prasanna · 2024
Whole slide image (WSI) classification in digital pathology is a challenging weakly supervised task due to the gigapixel scale of the data. While handcrafted features bring domain-specific insights, deep learned features offer superior generalizability and performance. Drawing inspiration from the attention mechanism in transformers, we introduce CoCa-MIL, a novel framework that unifies these features using Multiple Instance Learning (MIL). CoCa-MIL comprises two methods: Co-Attention, which leverages handcrafted features to guide deep feature-based representation learning, and Cross-Attention, which fuses both feature types to harness their complementary information for slide-level tasks. In this study, we show that both methods surpass traditional singlefeature-type WSI classification. On the TCGA Lung Cancer dataset, they achieve accuracy improvements of up to 2.60% and 5.21% over their respective baselines, underscoring the efficacy of attention-based fusion methods in exploiting the complementary nature of the handcrafted and deep features for enhancing performance beyond deep learning alone.