Embeddings de Tokens Iniciales de Modelos Basados en BERT para Identificar Texto Escrito por Humano o Generado Automáticamente

César Espín-Riofrio, Luis Ramos-Ramírez, Holger Camacho-Villalva, Débora K. Preciado-Maila, Jorge L. Charco, Arturo Montejo‐Ráez · 2024

Remarkable advances in text generation models have significantly expanded their applicability in a wide variety of fields.It is difficult to identify whether a text has been written by human or automatically generated, due to the ability of these models to mimic human style, coherence and expression.In this research, a Deep Learning method focused on Natural Language Processing (NLP) is proposed to identify the origin of a text.It is based on the extraction of the embeddings of the initial tokens of the twelve hidden layers of BERT-based Transformers models.The dataset provided in the IberLEF 2023 AuTexTification task was used, with texts extracted from different domains, in English and Spanish language.The DeBERTa model was used for the English texts and mDeBERTa for the Spanish texts.Optuna was used to automate the search for the optimal hyperparameters for the final training, performing fine-tuning of each model for its subsequent prediction and evaluation.The evaluation results of the proposed model were excellent, while the prediction results were not so good, being an interesting point for discussion and analysis of the proposal.

Read the paper · More papers on PaperTik