A Framework for Enhancing Accuracy in AI Generated Text Detection Using Ensemble Modelling
Kush Aggarwal, Sahib T Singh, Parul, Vipin Pal, Satyendra Singh Yadav · 2024
With the proliferation of AI-driven technologies, the generation of synthetic text has become increasingly prevalent, posing significant challenges in distinguishing between human-generated and AI-generated content (LLM). To mitigate this challenge, novel approach is proposed in this paper for AI-generated text detection through ensemble modelling based framework, leveraging the strengths of multiple state-of-the-art language models. Proposed Ensemble model integrates BERT, DeBERTa, and a custom ensemble method, each contributing to the collective decision-making process with weighted predictions. A diverse dataset sourced from various online platforms is used and this dataset comprises both human-written and AI-generated text samples. A fine-tuning strategy is used that dynamically adjusts the weights of the ensemble model based on the validation accuracy of each constituent model, while applying a cosine learning rate scheduler during training to optimize performance. The effectiveness of the ensemble model is evaluated using standard performance metrics such as accuracy, recall and F1 score. Proposed model achieved an accuracy over 94% and high recall of 98% through the ensemble framework, demonstrating accuracy improvement by 4.1% over BERT and and robustness in detecting AI-generated text across different domains and languages. The research contributes to advancing the field of AI-generated text detection and addresses critical challenges in content moderation and verification in online environments.