Contributions à la création de corpus et de modèles d’apprentissage profond pour les données textuelles multilingues

Anna Pappa · HAL (Le Centre pour la Communication Scientifique Directe) · 2024

This HDR synthesizes nearly a decade of research work in Natural Language Processing (NLP). It places a particular emphasis on the creation of multilingual and thematic corpora, specifically designed for sentiment and aspect analysis. The methodologies and tools developed for generating noiseless, multilingual datasets, sourced from user reviews, serve as a solid foundation for subsequent experiments. The use of hierarchicalConvolutional Neural Networks CNN and Recurrent Neural Networks (RNN) addresses the challenge of polarity prediction and thematic classification. This performance is further enhanced by the Bi-CNN-LSTM architecture, which combines convolutions with long-term memory, achieving an accuracy ranging from 90% to 100% depending on the experiments, and this on non-annotated multilingual corpora. Integrated deep learning techniques, such as transfer learning and active learning within a combined Bi-LSTM-CNN-CRF architecture, are employed for aspect annotation, thus improving the model’s performance, especially in contexts where data or languages are underrepresented. In summary, this habilitation contributes to the methods and practices in NLP, relying on tailor-made datasets and sophisticated model architectures to overcome complex challenges in semantic annotation and multilingual analysis.

Read the paper · More papers on PaperTik