Extracting Chinese Frequent Strings Without Dictionary From a Chinese corpus, its Applications.

Yih-Jeng Lin, Ming‐Shing Yu · Journal of information science and engineering · 2001

This paper describes how to extract Chinese frequent strings without using a dictionary. In this paper, we generalize the notations of words and unknown words to those of frequent strings. The Chinese frequent strings (CFSs) we define include words, unknown words, and other strings that are frequently used. Some examples of CFSs are “ (can only let)”, “ (every minute and every second)”, “ (bearing in mind the interest of each other)”, and “ (and nobody)”. A CFS is very useful in Chinese natural language processing and its related applications. We show its application to the following three tasks: Chinese phoneme-to-character conversion, Chinese character-to-phoneme conversion, and the determination of prosodic segments in a Chinese sentence for text-to-speech output. We have also developed a simple method to extract CFSs from a corpus. The method we propose can automatically detect such strings without the use of any lexicon, and no word segmentation is needed. We also can extract unknown words in a corpus which consist of three of more words. Such words (e.g. ) usually cannot be extracted by using a concatenation approach.

Read the paper · More papers on PaperTik