An Ensemble LLM Framework of Text Recognition Based on BERT and BPE Tokenization

Zien Huang · 2024

Given how quickly artificial intelligence is developing, Large Language Models (LLMs) such as GPT and BERT have achieved significant performance in text generation and language understanding. These models, trained on massive data sets, can now highly imitative human writing styles and the logic of human thinking, making the text they generate difficult to distinguish from human-written text in certain contexts. As the text generation capabilities of LLMs become increasingly powerful, the text automatically generated by machines may lead to misinformation and the spread of false content. The ability to discriminate between writing created by machines and text written by humans is becoming more and more crucial in both academics and industry. Developing effective identification methods for precise judgment of text has significant practical significance. Therefore, this study aims to explore a method to differentiate between human-written and LLM-generated texts. This paper discusses an integrated framework that uses the Byte-Pair Encoding (BPE) tokenization method to segment text, separately trains and constructs a deep learning model based on BERT fine-tuning, and a traditional machine learning model, combined with ensemble learning algorithms, to train an efficient classifier. The purpose of this classifier is to identify whether a piece of text is generated by a machine. The final model performs excellently on the test sets. This research is not only of great significance to academic research but also has practical application value in the authenticity identification of texts in news media, the publishing industry, and even at the legal level.

Read the paper · More papers on PaperTik