Automated Timeline Extraction via a Semi-supervised Learning Approach

Oulin Yang · 2011

This thesis relates to a new developing topic in the Natural Language Processing area of recent years extracting time and events from news articles. I developed a Java artefact, which is trained and evaluated on tagged news corpus via both supervised and semi-supervised learning methods. Specifically, this is done first through four baseline algorithms: a ‘Dictionary Lookup’ classifier, variants of Naive Bayesmodel, the TARSQI toolkit and the Conditional Random Fields (CRFs) classifier. The experimental results show that, by applying tokenization, part-of-speech tagging, preceding adjacent labels and other features, one variant of Naive Bayes algorithm gives better performance than other variants. Moreover, the CRF classifier outperforms other baselines significantly and beats the ‘stateof-the-art’ toolkit. To address the bottleneck brought by the limited size of training data set, this thesis introduces a novel framework of applying co-training algorithm on various combinations of CRFs with distinct feature sets, to automatically generate high quality labelling outputs. The labelled data set produced from co-training’s peak performance has been experimented with baseline classifiers. As a result, baselines are performing stronger with the new training data set than with the original data set. Overall, the major contribution of this thesis is finding an applicable way to use cotraining algorithm to process sequential text data, and promote general classification performance over unlabelled data. By this method, scarcity of training data in many circumstances can be eased to some extent.

Read the paper · More papers on PaperTik