Enabling Annotation of Historical Corpora in an Asynchronous Collaborative Environment

Enrique Manjavacas Arévalo, Peter Petré · 2017

Current research in Corpus Linguistics and related disciplines within the multi-disciplinary field of Digital Humanities, involves computer-aided manual processing of large text corpora. Typically, corpus instances are retrieved with the help of concordancers and textual search engines and subsequently labeled by hand before being submitted to quantitative analysis. While well-established software solutions already exist for corpus data retrieval, less attention has been paid to the annotation process in terms of both software facilities and best practices, especially in the context of collaborative research. However, with the increase in size and scope of research projects we envisage new needs for synchronizing interdependent annotations by different researchers. Current ad-hoc solutions to collaborative corpus analysis and annotation typically involve general-purpose Real-Time Editing (RTE) and cloud storage software, whose functionality is arguably sub-optimal for research purposes. In the present paper we discuss potential problems related to synchronizing annotations in large-scale projects, as well as the potential benefits that can be derived from a dedicated approach to annotation data management. As a proof of concept, we showcase our current solution in the form of Cosycat (Collaborative Synchronized Corpus Annotation Tool), a collaborative asynchronous application that has grown out of a Historical Linguistics research project involving several parallel studies and multiple researchers.

Read the paper · More papers on PaperTik