Focused training sets to reduce noise in NER feature models

Amber McKenzie · 2013

Feature and context aggregation play a large role in current NER systems, allowing significant opportunities for research into op-timizing these features to cater to different domains. This work strives to reduce the noise introduced into aggregated features from dis-parate and generic training data in order to al-low for contextual features that more closely model the entities in the target data. The pro-posed approach trains models based on only a part of the training set that is more similar to the target domain. To this end, models are trained for an existing NER system using the top documents from the training set that are similar to the target document in order to demonstrate that this technique can be applied to improve any pre-built NER system. Initial results show an improvement over the Univer-sity of Illinois NE tagger with a weighted av-erage F1 score of 91.67 compared to the Illinois tagger’s score of 91.32. This research serves as a proof-of-concept for future planned work to cluster the training docu-ments to produce a number of more focused models from a given training set, thereby re-ducing noise and extracting a more repre-sentative feature set. 1

Read the paper · More papers on PaperTik