Boundary Analysis in Ensemble Clustering for Improved Document Classification

Ali Alsalama, Ashraf Elnagar · 2025

We propose a novel ensemble clustering approach for multi-label document classification by decoding cluster boundaries. Our methodology leverages the strengths of multiple clustering algorithms, including Fuzzy C-Means (FCM), Gaussian Mixture Model (GMM), HDBSCAN, K-Means, and BIRCH, to achieve robust document clustering. We enhance classification accuracy by integrating cluster outputs through ensemble learning, leveraging AraBERT embeddings for improved representation of Arabic text. Using the SANAD dataset of Arabic news articles, we preprocess the data by removing temporal biases and noise, then generate contextual embeddings with AraBERT. A hybrid strategy is employed, where single-label classification is performed via majority-vote ensembles, while multi-label classification assigns documents to multiple clusters based on confidence scores. This approach captures the nuanced and overlapping nature of document topics, achieving an overall clustering accuracy of 97.8% and significantly improving classification robustness.

Read the paper · More papers on PaperTik