A Hybrid Machine Learning Approach for Document Classification: A Comparative Study
Thota Balaji, V Khanna, T. Nalini · 2023
In the realm of Natural Language Processing (NLP), document categorization is a difficult problem since it requires assigning documents to preset groups based on their content. In this study, a novel method for document categorization that combines K-Means Clustering (KMC) with the Latent Dirichlet Allocation (LDA) algorithm was provided. Documents are first clustered using K-means clustering, and then latent topic modeling is used in each cluster to extract latent themes. Documents are sorted into their appropriate categories using the extracted subjects. The suggested technique on two open-source datasets and comparing the outcomes to those obtained by existing state-of-the-art approaches were tested. The outcomes from experiments proved the superior accuracy, precision, and recall of the suggested strategy over state-of-the-art approaches.