Docria: Processing and Storing Linguistic Data with Wikipedia
Marcus Klang, Pierre Nugues · DSpace repository (University of Tartu) · 2019
The availability of user-generated content has increased significantly over time.Wikipedia is one example of a corpus, which spans a huge range of topics and is freely available.Storing and processing such corpora requires flexible document models as they may contain malicious or incorrect data.Docria is a library which attempts to address this issue with a model using typed property hypergraphs.Docria can be used with small to large corpora, from laptops using Python interactively in a Jupyter notebook to clusters running mapreduce frameworks with optimized compiled code.