Enhancing Digital Forensics: Machine Learning Techniques for Social Media Investigation

Dhairya Vyas, Milind Shah, Ankita Kothari, Jyoti Golakia, Vruti Parikh · Procedia Computer Science · 2025

Due to the increasing growth of social media platforms, improved methods are needed to extract, examine, and interpret digital evidence. Due to the vastness and ever-changing nature of the data that is collected via social media platforms, traditional forensic procedures usually meet challenges. Specifically created for the purpose of conducting forensic analysis of material obtained from social media platforms, the purpose of this research is to build and analyze machine learning models and algorithms. Gathering information from various platforms, such as Kaggle, is the first step in the process. The succeeding processes involve data cleaning and categorization, which includes activities such as text normalization, entity recognition, and sentiment analysis. These steps come after the first data gathering has been completed. The extraction of relevant information, such as user behaviors, temporal trends, and network connections, is accomplished through the use of feature engineering. A significant portion of the investigation is centered on the application of supervised machine learning techniques for the purpose of detecting users and detecting offensive language. In order to determine whether or not the proposed approach is effective, comprehensive experiments are conducted utilizing datasets derived from social media activities. The findings indicate that there have been considerable improvements in the capabilities of forensic investigations, including enhanced accuracy, scalability, and automation. In this research we have investigated social media’s data and detected offensive language using various machine learning algorithms such as Support Vector Machine (SVM), K-Nearest Neighbours (KNN), Decision Tree, Random Forest, and Extra Tree. The support vector machine (SVM) achieved 86% accuracy where as KNN achieved only 77% accuracy which was far lower than other learning algorithms. However, the Decision Tree, Random Forest, and Extra Tree models achieved 99% accuracy with good precision, recall, and F1-scores, proving their ability to detect and categorize social media information.

Read the paper · More papers on PaperTik