Improved Text Categorisation for Wikipedia Named Entities
Sam Tardif, James Curran, Tara Murphy · 2009
The accuracy of named entity recognition systems relies heavily upon the volume and quality of available training data. Improving the process of automatically producing such training data is an important task, as manual acquisition is both time consuming and expensive. We explore the use of a variety of machine learning algorithms for categorising Wikipedia articles, an initial step in producing the named entity training data. We were able to achieve a categorisation accuracy of 95 % F-score over six coarse categories, an improvement of up to 5 % F-score over previous methods. 1