Unknown Word and Phrase Extraction Using a Phrase-Like-Unit-Based Likelihood Ratio

Yu‐Sheng Lai, Chung‐Hsien Wu · International Journal of Computer Processing Of Languages · 2000

In this paper, we propose a statistical method to extract unknown words and phrases from sentences in a specific domain. The unknown words or phrases are defined as phrase-like units (PLU) that can be the combinations of some words in the lexicon or some characters that appear together. A PLU-based likelihood ratio is proposed to extract possible PLUs. Two principles, overlap competition and inclusion competition, are used to decide the final unknown words or phrases. The method not only can detect unknown words or phrases but also can correct the errors from word segmentation. Additionally, this method is expandable and portable to other domains. We collected articles from MSDN (Min Sheng Daily News) over half a year to construct a corpus containing over 275,000 sentences in the corpus. We randomly choose 175,000 sentences from the corpus as the experimental corpus. The recall rate and the precision rate are used to evaluate our proposed method. Using the extraction method, the recall rate achieved 88.7% while the precision rate achieved 88.2%. The experimental results show that the method can detect unknown words and phrases automatically and efficiently from sentences in a specific domain.

Read the paper · More papers on PaperTik