Presenting the Bangor Autoglosser and the Bangor Automated Clause-Splitter
D M Carter, Mirjam Broersma, K. Donnelly, Agnieszka E. Konopka · Digital Scholarship in the Humanities · 2017
Until recently, corpus studies of natural bilingual speech and, more specifically, codeswitching in bilingual speech have used a manual method of glossing, part-of-speech tagging, and clause-splitting to prepare the data for analysis. In our article, we present innovative tools developed for the first large-scale corpus study of codeswitching triggered by cognates. A study of this size was only possible due to the automation of several steps, such as morpheme-by-morpheme glossing, splitting complex clauses into simple clauses, and the analysis of internal and external codeswitching through the use of database tables, algorithms, and a scripting language.