Surplus Data Prediction and Classification of Textual-Data Using Machine and Deep Learning Comparative Analysis

N. R. Rajalakshmi, S. Saravanan, Anupam Singha · 2023

In recent years, several classification techniques have become available to categorize texts efficiently. In the field of machine learning, the construction of classifiers involves analyzing and learning from well-labeled datasets with specified categories. Likewise, deep learning's performance improvement in the text categorization sector comes from its ability to achieve high precision with a less complex setup and processing. This work emphasizes machine learning, and deep learning techniques are used to classify textual data. On account of the fact that textual data often carries a significant amount of unnecessary data, pre-processing is a crucial step, e.g.,. fill in the blanks and remove duplicate rows, then clean up the data. Following this, deep learning algorithms such as bidirectional long shortterm memory (BiLSTM), recurrent weighted average (RWA), and conditional random field (CRF) are used for classification, along with traditional machine learning algorithms like naive Bayes, gradient boosting, and support vector machines (SVMs). This work reveals that BiLSTM turned out to be the most accurate, achieving a classification accuracy of 98.5 %, outperforming both other models and baseline investigations.

Read the paper · More papers on PaperTik