AI-based Framework for Discriminating Human-authored and AI-generated Text
Mudasir Ahmad Wani, Ahmed A. Abd El‐Latif, Mohammad ELAffendi, Amir Hussain · IEEE Transactions on Artificial Intelligence · 2024
Deep learning techniques are increasingly adept at distinguishing between human-written and AI-generated text. This study presents a deep learning-based text classification system to discern the source of text—whether it’s generated by AI or authored by a human. We compiled a comprehensive and diverse dataset incorporating existing datasets, real-time tweets, and synthetic data generated by cutting-edge AI models such as Google Bard and OpenAI’s ChatGPT. Seven deep learning techniques including Convolutional Neural Networks (CNN), Artificial Neural Networks (ANN), Recurrent Neural Networks (RNN), Long Short-Term Memory networks (LSTM), Gated Recurrent Unit (GRU), Bidirectional LSTM (BiLSTM), and Bidirectional RNN (BiRNN) and four text embedding techniques namely One-hot, Word2Vec, TF-IDF, and BERT, are employed to perform 28 experiments (4*7). Our findings exhibit outstanding performance across all models. BERT-based models notably outshine the others, with BERT-CNN achieving the highest accuracy of 98.87%, an MCC of 0.9742, and an F1 measure of 0.9916. Following closely, Bert-BiLSTM attained an accuracy, MCC, and F1 score of 98.77%, 0.9720, and 0.9909, respectively.This paper demonstrates the effectiveness of deep learning approaches in distinguishing between human and AI-generated text, with potential applications in automated content moderation, detection of AI-generated spam, and safeguarding the authenticity of user-generated content.