Computational workshop: N-Gram processor

Andreas Buerki · ORCA Online Research @Cardiff (Cardiff University) · 2017

In this workshop, participants are taken through the process of extracting items of formulaic language from a set of corpus documents, assessing the quality of the extraction and using the resulting list to annotate text documents for formulaicity. This is done using the N-Gram Processor (Buerki 2013) and SubString (Buerki 2011) software packages and Wikipedia text files.

Read the paper · More papers on PaperTik