Computational workshop: N-Gram processor
Andreas Buerki · ORCA Online Research @Cardiff (Cardiff University) · 2017
In this workshop, participants are taken through the process of extracting items of formulaic language from a set of corpus documents, assessing the quality of the extraction and using the resulting list to annotate text documents for formulaicity. This is done using the N-Gram Processor (Buerki 2013) and SubString (Buerki 2011) software packages and Wikipedia text files.