Word frequency distribution in Japanese text*

Koichi Ejiri, Niklaus Staeheli, Shiori Ooaku · Journal of Quantitative Linguistics · 1994

In our last paper (Ejiri 1993), we proposed a new parameter, G = log (N/L)/{log(N)‐1}, where N is the number of words and L the number of different words, which correlates with constraint of the target text represented in ASCII code. We found that the same measure is applicable to Japanese texts which have no clear word segmentation. By statistical analysis of kanji, kata‐kana and alphabetic strings, Japanese texts were found to have a similar distribution as English texts or computer language. We also introduced a “joint entropy” of string[i] and string[j] where the latter string follows the former string after a fixed distance. Here, the distance means the number of words (or defined character strings) between string[i] and string[j]. This entropy is a measure of redundant description (phrases) in the text. A simpler approach using n‐gram frequency was found to be useful to detect errors in a text recognized by OCR (Optical Character Reader).

Read the paper · More papers on PaperTik