Impact of Preprocessing on Twitter Based Covid-19 Vaccination Text Data by Classification Techniques

P.C. Sridevi, Thambusamy Velmurugan · 2022 International Conference on Applied Artificial Intelligence and Computing (ICAAIC) · 2022

Data from sources like twitter, organizations internal communications, Face book, blogs are keys in providing user the choice and scrutiny in organizational elevating. Text classification is used unrestrictedly to assign a fixed predefined grouping. This grouping is used for articles categorizing, chat organizing and brand mention on people opinion. There are number of method to classify, cluster or associate (Sentiment analysis activities) these data so as to obtain public opinion regarding particular topic. To perform sentiment analysis activities of the data collected from the source need to be processed. Garbage-in garbage-out is pertinent in relationships to data analysis. That is the output the analyzer gets from the analysis is completely depend on the input data provided. The inconsistent, jagged, inadequate, erroneous, data produce inaccurate results even though analyzer uses powerful algorithm. A preprocessing step is necessary to convert raw dirty data to trainable, understandable and analyzing data in the sentiment analysis algorithm. In this research work, twitter COVID data set is classified using LIBLINEAR and BayesNet classification techniques on both processed and unprocessed data. For processing, resample and Remove Useless filter is used and the results are compared. For the comparison purpose performance metrics like Fl-score and precision were also used to validate the used models with its results. Through this comparison analysis, the best performances of the dassification algorithms are suggested for further process.

Read the paper · More papers on PaperTik