Multimodal pretraining for histology-gene joint representation learning

Danial Maleki, Nazim Shaikh, Gareth Shannon, Jian Li, Yao Nie · 2025

Computational pathology has shown great promise in developing prognostic models from histology images. Recent advancements in multimodal models have demonstrated that integrating whole-slide images and bulk transcriptomics data can further improve patient prognosis understanding and prediction. However, most existing methods require both modalities as the inputs to conduct the prediction, hindering the technology adoption in practical applications due to the limited availability of high cost modalities. Our work addresses this challenge by enhancing a multimodal pre-training framework to learn histology-gene joint representations, which incorporate both morphology and molecular information, while being generated using image data as input alone. Specifically, an advanced vision transformer model pre-trained on pathology images was utilized to generate tissue image embeddings to facilitate the multimodal pre-training. This could result in more robust joint representations and better performance in patient survival prediction for both colorectal and lung cancer patients. Additionally, the scheme might be extended to address the gene mutation prediction task for the lung cancer patients, potentially showing improvement compared to the approach without multimodal pre-training.

Read the paper · More papers on PaperTik