Transformer Based News Text Classification

Junshuai Li · Applied and Computational Engineering · 2025

With the exponential growth of online news, Transformer models based on self-attention mechanisms (e.g., BERT, GPT) have demonstrated theoretical advantages over traditional methods (e.g., SVM, Naïve Bayes, and CNN) in news text classification by capturing global semantic relationships. The encoder-only Transformer architecture developed in this study, integrating multi-head self-attention, dynamic positional encoding, and global average pooling, achieved an initial accuracy of 69.52% on the 20 Newsgroups dataset (significantly higher than CNN's 57.59%), showcasing its superior global feature extraction, adaptive polysemy handling, and noise resilience. However, the model suffers from prolonged training times (1,252 seconds per epoch compared to CNN's 149 seconds) and late-stage overfitting. Despite computational efficiency challenges, the research proposes optimizing performance through sparse attention mechanisms, domain-specific pretraining, and hybrid Transformer-CNN architectures to enhance classification capabilities in long-text and multilingual scenarios. These findings validate Transformer's potential for complex NLP tasks while emphasizing the necessity of architectural refinements to balance performance and scalability, providing critical directions for advancing news classification systems.

Read the paper · More papers on PaperTik