A Method for AI-generated sentence detection through Large Language Models

Fabio Martinelli, Francesco Mercaldo, Luca Petrillo, Antonella Santone · Procedia Computer Science · 2024

In recent years, we have seen an impressive expansion of a family of Artificial intelligence models, known as generative AI, that are capable of producing fresh, unique material, including text, images, audio, and even code. Large datasets of previously published information are used to train these models, enabling them to mimic the patterns and structures of the data and produce original output that is stylistically and qualitatively comparable to the training set. While these models have many promising applications, they also carry significant risks and potential dangers that must be carefully considered, such as misinformation, intellectual property violations, or biased information. For these reasons, in this work, we have proposed a method to detect whether a sentence is human-generated or AI-generated. To achieve this goal, we used a labeled dataset to train four different models from the BERT family, achieving an Accuracy of 96%.

Read the paper · More papers on PaperTik