Application of BERT Model for Unsupervised Text Classification using Hierarchical Clustering for Automatic Classification of Thesis Manuscript

Cristian D. Montes, Junisare V. Silvosa, Cristopher C. Abalorio, Regien B. Nakazato · 2024

Many current studies on natural language processing (NLP) depend on supervised learning, which needs a lot of labeled data. This is not practical for large classification tasks. This research presents an unsupervised approach to automatically classify unlabeled theses using a BERT-hierarchical model. This technique combines BERT, an open-source machine-learning tool for NLP, with divisive hierarchical clustering utilizing the K-means clustering algorithm. Standard metrics evaluated the model's performance. Fine-tuning the BERT-hierarchical model with a learning rate of 0.00001 gave the best results, achieving 93% accuracy, 91% precision, and 93% recall. Additionally, a web application was created to facilitate model testing, offering functions for data creation, editing, and deletion. This study shows the BERT-hierarchical model's effectiveness in classifying unlabeled data, providing a scalable solution for large NLP tasks.

Read the paper · More papers on PaperTik