Improvement of Named Entity Tagging by Machine Learning

Thamar Solorio, Luís Enrique Erro · 2004

A Named Entity (NE) is a word, or sequence of words that can be classified as a name of a person, organization, location, date, time, percentage or quantity. Named entities can be valuable in several natural language applications. For instance, automatic text summarization systems can be enriched by using NEs, as they provide important cues for identifying relevant segments in text. Other uses of NE taggers are in the fields of information retrieval (i.e. more accurate Internet search engines), information extraction, automatic speech recognition, question answering and machine translation. There has been a considerable amount of work that aims to improve the performance of NE taggers, however most of these efforts are targeted to build handcrafted NE taggers. While handcrafted systems can achieve good performance they have several disadvantages, such as the need of rebuilding the NE tagger in order to port it to new domains. This proposal presents a research project for building an automated method, based on machine learning, that facilitates the portability of handcrafted NE taggers. The hypothesis underlying this proposal is that the coverage of NE taggers can be increased using machine learning techniques, without the need of rebuilding the linguistic resources of the NE tagger. One motivation for this research is that the task of building a training set is faster and easier than building the linguistic resources for the handcrafted systems. Preliminary results presented here show that it is feasible to prove the correctness of the hypothesis.

Read the paper · More papers on PaperTik