Improved Text Categorisation for Wikipedia Named Entities

Sam Tardif, James Curran, Tara Murphy · 2009

The accuracy of named entity recognition systems relies heavily upon the volume and quality of available training data. Improving the process of automatically producing such training data is an important task, as manual acquisition is both time consuming and expensive. We explore the use of a variety of machine learning algorithms for categorising Wikipedia articles, an initial step in producing the named entity training data. We were able to achieve a categorisation accuracy of 95 % F-score over six coarse categories, an improvement of up to 5 % F-score over previous methods. 1

Read the paper · More papers on PaperTik