Algorithmic Application and Research on BERT Model in English Topic Categorization

Tingting Chen · 2025

Since traditional text categorization methods often have unsatisfactory model performance in environments with complex topic semantics and sparse data. In this paper, a BERTopic English topic classification method based on BERT embedding is proposed. Firstly, the accuracy and adaptability of topic classification are improved by combining BERT embedding, UMAP downgrading and HDBSCAN clustering. Subsequently, the c-TF-IDF technique is used to optimize the topic word extraction and improve the interpretability of the model. The experimental data are derived from IEEE published literature metadata and rigorously preprocessed to construct a high-quality topic corpus. The results show that BERTopic significantly outperforms traditional topic modeling approaches (LDA and Top2Vec) in terms of accuracy, F1 score, precision and recall, especially when dealing with semantic complexity and category imbalance. Further analysis shows that the bidirectional contextual semantic representation capability of BERT embeddings can better capture the deep semantic relationships in text, especially on large-scale datasets with significant advantages. This study provides an efficient and robust solution for topic categorization tasks in a variety of scenarios such as information retrieval, content clustering, and document analysis. Future work could explore the potential of the model for application on multilingual and multimodal datasets.

Read the paper · More papers on PaperTik