How Natural Language Processing Enables AIGC Recognition? --Latest Trends and Future Prospects

Weizheng Wang, Hong Qiao · 2024

As the technology behind large language models advances rapidly, AI-generated content (AIGC) pervades our daily lives. Classifiers that identify AIGC play a crucial role in distinguishing between text generated by humans and that generated by artificial intelligence. In order to better prevent the abuse of AIGC and reduce the emergence of issues such as false information, academic misconduct, and deceptive comments, we introduced the task of AIGC classifiers, emphasizing the necessity of classifier development in this era. The essence of AIGC identification tasks lies in binary classification, aiming to discern whether a piece of content is created by artificial intelligence. In recent years, white-box and black-box methods as classifiers for identifying AIGC have made significant strides. In this paper, we curated the main research achievements in the field of AIGC identification, emphasizing the crucial role of comprehensive and excellent datasets in constructing AIGC recognition classifiers. Additionally, we explored the limitations and development goals of current popular datasets, as well as potential datasets. Furthermore, we analyzed paradigms of various classifiers, addressing challenges such as multidomain recognition tasks, cross-language recognition tasks, and data ambiguity issues. Finally, we proposed pathways for the future development of AIGC identification. This study aims to provide a clear overview for relevant researchers and offer constructive suggestions for constructing more stable and efficient classifiers.

Read the paper · More papers on PaperTik