Analysing Wikipedia and gold-standard corpora for NER training
Joel Nothman, Tara Murphy, James Curran · 2009
Named entity recognition (ner) for English typically involves one of three gold standards: muc, conll, or bbn, all created by costly manual annotation. Recent work has used Wikipedia to automatically create a massive corpus of named entity annotated text.