Enhancing AI-Generated Text Identification with BERT-CNN and DistilBERT-BiLSTM Models

Rajsekhar Das, Ricky Dey, Sorbojit Mondal, Nabanita Das, Bikash Sadhukhan · 2025

This study proposes two advanced transformer-based architectures for enhancing the identification of AI-generated text: a fine-tuned BERT-CNN model and a hybrid DistilBERT-BiLSTM framework. The BERT-CNN architecture combines pre-trained BERT embeddings with convolutional neural networks to detect localized linguistic patterns indicative of synthetic text. The DistilBERT-BiLSTM model integrates the efficiency of DistilBERT with bidirectional LSTM layers to capture sequential dependencies and long-range contextual features. Both approaches employ standardized preprocessing using the BERT tokenizer, including tokenization, padding, and truncation, to ensure consistency in input representation. The BERT-CNN model achieved strong performance with 95.67% accuracy, 94.32% F1-score, and 93.45% precision, demonstrating its capability to discern subtle AI-generated patterns. The DistilBERT-BiLSTM framework further enhanced detection accuracy to 97%, with precision, recall, and F1-score values of 98%, 97%, and 97%, respectively, attributed to its ability to model temporal relationships in text sequences. Both models exhibited robustness against paraphrasing-based evasion techniques, with DistilBERT-BiLSTM showing superior generalization due to its balanced architecture of lightweight language understanding and sequential analysis. This research underscores the efficacy of transformer-based hybrid models in advancing AI-generated text detection, offering scalable solutions for maintaining content authenticity in academic, professional, and digital platforms. The findings contribute to the development of reliable tools for ethical AI adoption and mitigation of misinformation risks.

Read the paper · More papers on PaperTik