Deploying a co-training algorithm to classify human-rights abuses

Ragini Gokhale, Maria Fasli · 2017

With the advent of the Internet and social media data, we have seen a considerable growth in the availability of data related to human right abuses. Human Rights and Non-Government organisations, whose purpose is to provide rehabilitation services and support to victims of torture, require powerful and appropriate methods to learn from and draw insights from this data. Current approaches usually require large amounts of expensive labelled data in order to make accurate predictions. This paper proposes an approach to applying a domain specific term extraction method using an ontology to train a semi-supervised co-training classifier that would classify the data collected into relevant categories. The method implements and utilises a domain ontology specifically created for the domain of human rights as background knowledge to extract the preliminary terms for generating the labelled data in order to train the classifier. Out experiments show that this method improves over the individual classifiers, while also displaying a significant reduction in the amount of labelled data required for training accurate classifiers.

Read the paper · More papers on PaperTik