Augmenting Document Classification Accuracy Through the Integration of Deep Contextual Embeddings
Rama Krishna Paladugu, Gangadhara Rao Kancherla · Ingénierie des systèmes d information · 2024
Document classification, a fundamental process within the field of natural language processing, has benefitted from the recent advancements in deep learning, particularly in enhancing accuracy.Traditional text clustering methods, such as bag-of-words models, exhibit domain specificity and struggle to handle vast data volumes.They also face limitations in elucidating sophisticated patterns and intricate word and phrase relationships within textual data.These constraints may adversely affect the accuracy of text clustering, subsequently impacting downstream applications like information retrieval, document classification, and natural language processing.This paper proposes a novel text classification model, termed Deep Contextual Embeddings Model (DCEM), designed to improve document classification accuracy.The DCEM integrates pre-trained deep contextual embedding architectures (e.g., GPT-2) with text clustering algorithms (e.g., K-Means).It employs contextual embedding models to enhance document clustering accuracy by capturing context and semantic depth, improving data structure comprehension, and eliminating noise for more precise results.Experimental results, derived from the application of DCEM on AG News, Reuters-21578, and IMDB reviews datasets, indicate a significant improvement in document classification accuracy (81.09%), compared to traditional text clustering and document classification methods.