Named Entity Extraction for Information Retrieval 1
Hsin‐Hsi Chen, Yung-Wei Ding, Shih-Chung Tsai · 1998
Name extraction is indispensable for both natural language understanding and information retrieval. However, proper names are major unknown words in natural language texts, and unknown word identification is still a challenge problem in natural language proces- sing. This paper deals with identification of person names, organization names and location names from Chinese texts. Different types of information from different levels of text are employed, including character conditions, statistic information, titles, punctuation marks, or- ganization and location keywords, speech-act and locative verbs, cache and n-gram model. We also clarify which strategies can be used in which cases, i.e., queries and/or documents. In our experiments, the recall rates and the precision rates for the extraction of person names, orga- nization names, and location names under MET data are (87.33%, 82.33%), (76.67%, 79.33%) and (77.00%, 82.00%), respectively.