Impact Analysis of Text Representation on Biomedical Multi-Label Text Classification with Deep Learning
Hemraj Kumawat, Aditi Sharan, Shikha Verma · Procedia Computer Science · 2025
Multi-label text classification is a challenging task compared to single-label text classification due to the need to predict multiple, non-exclusive labels simultaneously for a single text document. Deep learning models can capture such label dependencies. However, classification performance depends on the specific deep learning architecture and how the text is represented numerically for input to deep learning models. This research examines the impact of text representation on multi-label text classification tasks using various text representation techniques, such as traditional TF-IDF, embedding-based models like GloVe, Word2Vec, fastText, and transformer-based models like BERT, BioBERT, SciBERT, and PubMedBERT. This paper utilizes two publicly available datasets for multi-label text classification. The first dataset, the Chemical Exposure Information (CEI) Corpus, consists of a collection of PubMed abstracts labeled with one or more classes using an exposure taxonomy of 32 classes. The second dataset, Hallmarks of Cancer (HoC), consists of a collection of PubMed abstracts manually labeled with one or more Hallmarks of cancer classes. The performance of these models was evaluated using micro-precision, micro-recall, micro-F1 score, 0/1 subset accuracy, hamming loss, jaccard score, coverage, ranking loss, and label ranking precision. The results reveal that model performance varies significantly, with PubMedBERT, a transformer-based model, outperforming other models with micro-precision, micro-recall, and micro-F1 score values of 0.93, 0.89, and 0.91, respectively, for the CEI dataset and with micro-precision, micro-recall, and micro-F1 score values of 0.88, 0.78, and 0.83 respectively, for HoC dataset. This study emphasizes the necessity of employing suitable text representation techniques to increase the efficiency of text classification applications such as literature indexing and information retrieval.