Some Distributional Properties of Mandarin Chinese : A Study based on the Academia Sinica Corpus
Ching‐Yu Chen, Shu‐Fen Tseng, Chu‐Ren Huang, Keh-Jiann Chen · Institutional Repositories DataBase (IRDB) · 1993
The study of word frequency has been discussed by linguists, psychologists, and computer scientists.However, the results of these studies cannot be valid unless the corpus is big enough and properly-segmented.This paper observes the distributional information derived from word frequency based on .a 14-million-character corpus of Chinese newspaper (CKIP 1993).This is the first available Mandarin Chinese corpus of such magnitude.The word frequency count is obtained with an automatic-segmentation program with above 99% accuracy rate (Chen and Liu 1992).The count reflects some general phenomena of Chinese usage.For example, among the first thousand high frequency words, there are more bi-syllabic words than mono-syllabic words, attesting to the trend of bi-syllabicfication observed by many linguists.However, in general, the mono-syllabic function words occur more frequently than bi-syllabic words.In addition, the frequency of numerals is ranked according to their numeric order ('one' is higher than 'two', and 'two' is in turn higher than 'three ', etc.)This paper discusses the theoretical and applicational implications of these distributional properties.For instance, we find that the most frequent 2452 characters and 28124 words make up 99% of the corpus content.It is suggested that the optimal strategy for learning Chinese lies in the mastery of the most frequent 2452 characters plus words whose meanings can not be predicted on the basis of their component characters.This implies that one need not know 28124 words in order to achieve good reading knowledge in Chinese.Given the noted parallel between the internal structure of words and phrases, one can predict that knowledge of a few thousand words and of the morphosyntactic rules will enable one to read 'Chinese without much difficulty.