Social Media Forensics :Cyberbullying Tweet Classification Using Natural Language Processing and Machine Learning
2025
Using machine learning and natural language processing techniques, the Cyberbullying Tweet Classification project presents an intelligent method for identifying and classifying damaging social media information.By grouping tweets into six categories Age, Ethnicity, Gender Religion Other Cyberbullying and Not Cyberbullying.This approach seeks to detect different types of cyberbullying in tweets.The objective is to give parents, digital forensics experts, and internet platforms a support network to identify abusive activity and advance a safer online environment The system uses Natural Language Processing (NLP) techniques, such as tokenization, stopword removal, punctuation cleaning, stemming, and lemmatization, to preprocess raw twitter text.TF-IDF (Term Frequency-Inverse Document Frequency) is then used to convert the cleaned texts into numerical feature vectors.Using this data, a machine learning model more precisely, a Support Vector Machine (SVM) with a linear kernel is trained to identify trends related to every kind of cyberbullying.According to evaluations, the model's accuracy is around 82.9.A web application built on Streamlit has also been created to enable real time user interaction with the system.Users can get instant predictions and a graphic depiction of the identified category by inputting a tweet into the app.The pickle module preserves the model and vectorizer for effective runtime loading.This study provides as a basis for developing scalable, multilingual, and real-time cyber safety solutions and illustrate the usefulness of natural language processing and machine learning in identifying online abuse.