Team MGTD4ADL at SemEval-2024 Task 8: Leveraging (Sentence) Transformer Models with Contrastive Learning for Identifying Machine-Generated Text

Huixin Chen, Jan Büssing, David Rügamer, Ercong Nie · 2024

This paper outlines our approach to SemEval-2024 Task 8 (Subtask B), which focuses on discerning machine-generated text from humanwritten content, while also identifying the text sources, i.e., from which Large Language Model (LLM) the target text is generated.Our detection system uses Transformer-based techniques and incorporates various pre-trained language models (PLMs), which are tools that help understand and process language, including sentence transformer models.Additionally, we incorporate Contrastive Learning (CL) into the classifier to improve the detecting capabilities and employ Data Augmentation methods.Ultimately, our system achieves a peak accuracy of 76.96% on the test set of the competition, configured using a sentence transformer model integrated with CL methodology.

Read the paper · More papers on PaperTik