The CoMeRe corpus for French : structuring and annotating heterogeneous CMC genres

Thierry Chanier, Céline Poudat, Benoît Sagot, Georges Antoniadis, Ciara R. Wigham, Linda Hriba, Julien Longhi, Djamé Seddah · LDV-Forum/Journal for language technology and computational linguistics · 2014

The CoMeRe project aims to build a kernel corpus of different computer-mediated communication (CMC) genres with interactions in French as the main language, by assembling interactions stemming from networks such as the Internet or telecommunications, as well as mono and multimodal, and synchronous and asynchronous communications.Corpora are assembled using a standard, thanks to the Text Encoding Initiative (TEI) format.This implies extending, through a European endeavor, the TEI model of text, in order to encompass the richest and the more complex CMC genres.This paper presents the Interaction Space model.We explain how this model has been encoded within the TEI corpus header and body.The model is then instantiated through the first four corpora we have processed: three corpora where interactions occurred in single-modality environments (text chat, or SMS systems) and a fourth corpus where text chat, email, and forum modalities were used simultaneously.The CoMeRe project has two main research perspectives: discourse analysis, only alluded to in this paper, and the linguistic study of idiolects occurring in different CMC genres.As natural language processing (NLP) algorithms are an indispensable prerequisite for such research, we present our motivations for applying an automatic annotation process to the CoMeRe corpora.Our wish to guarantee generic annotations meant we did not consider any processing beyond morphosyntactic labelling, but prioritized the automatic annotation of any freely variant elements within the corpora.We then turn to decisions made concerning which annotations to make for which units and describe the processing pipeline for adding these.All CoMeRe corpora are verified thanks to a multi-stage quality control process that is designed to allow corpora to move from one project phase to the next.Public release of the CoMeRe corpora is a short-term goal: corpora will be integrated into the forthcoming French National Reference Corpus, and disseminated through the national linguistic infrastructure Open Resources and Tools for Language (ORTOLANG).We, therefore, highlight issues and decisions made concerning the OpenData perspective.1 Introduction: the CoMeRe project Various national reference corpora have been successfully developed and made available over the past few decades, e.g. the British National Corpus (Aston and Burnard 1998), the SoNaR Reference Corpus of Contemporary Written Dutch (Oostdijk et al. 2008), the DWDS Corpus for the German Language of the 20th century (Geyken 2007), the DeReKo German Reference Corpus (Kupietz and Keibel 2009) and the Russian Reference Corpus (Sharoff 2006).Despite being in strong demand, no French national reference corpus currently exists.Thus, the Institut de la Langue Française (ILF) has recently taken the first steps to lay the groundwork for such a project.The aim is for the national project to both collect existing 1 CoMeRe stands for "Communication Médiée par les Réseaux," an updated equivalent to Computer-Mediated Communication (CMC) or Network-mediated Communication.

Read the paper · More papers on PaperTik