New York City Street Cleanliness: Apply Text Mining Techniques to Social Media Information.
Huijue Kelly Duan · 2020
Relevancy Determination• Apply a list of keywords to filter out some irrelevant tweets • Preprocess the data by applying tokenization and lemmatization, removing hashtags, @, URL links, stopwords, etc.• Use supervised machine learning models to identify relevant tweets, including Naïve Bayes, Random Forest, XG Boost o The dataset is facing an imbalanced classification issue, two sampling methods are used to solve the issue: Random under-sampling & Random over-sampling o Stratified 10-fold cross-validation with a paired t-test are performed to evaluate the performance of all classifiers (based on Accuracy, Precision, Recall, F-1 Score, ROC_AUC)