Linguistic Analysis of Tweets – Using Data Mining to study usage of English on Twitter
Puru Malhotra, Yugam Bajaj · 2020
The recent evolution in the language used by people in their day to day life has shown a major shift in trend in how they like to express themselves in public. With the increased use of abbreviations, shorthand notations, mnemonics and an introduction of a new set of hybrid words, the vocabulary used in communications happening between people has started to deviate from the original vocabulary that is traditionally used for a language. In this study, we have analyzed the data collected through a set of Tweets on Twitter to observe the shift of used vocabulary from the traditional vocabulary. Twitter is a popular social media platform where a variety of people express themselves openly and this makes it a good choice to extract samples of data without any bias. The tweets were extracted and analyzed by data mining techniques using the functions of the Twitter API for python. Each word was categorized into one of the two categories: words part of the traditional vocabulary of a language, what we are calling here as Proper English Words, and words not part of the traditional vocabulary, The Improper English Words, and then analysis was done over the observations. For this study, we have restricted ourselves to the tweets made in English only. The British English spellings were preferred.