Deepfake Text Detection Using NLP and ML

Akashdeep, Harshika Mitra, Ovais Bashir Gashroo · 2025

With the growing sophistication of large language models, distinguishing between human-written and machinegenerated text has become increasingly challenging. Deepfake text—content that mimics human writing but is produced by AI—poses serious risks in areas such as academia, journalism, and public discourse. This study presents a comprehensive approach to detecting deepfake text using Natural Language Processing and machine learning methods. We introduce a balanced, multilingual dataset consisting of English and Hindi text samples, both AI-and human-generated. After preprocessing and feature extraction using TF-IDF, we train and evaluate multiple machine learning models including SVM, XGBoost, Logistic Regression, and more. Our results show that traditional models, when properly tuned, can achieve exceptional performance—upto 99 for this task. The evidence endorses reliable scalable solutions for deepfake detection in pragmatic, multilingual contexts with the potential for even greater uptake across academia, media and cybersecurity for improved information authenticity verification. The research also offers an effective basis for future developments and encourages ongoing investigation of multilingual and cross-domain detection methods.

Read the paper · More papers on PaperTik