A computer readability formula of Japanese texts for machine scoring
Tateisi Yuka, Yoshihiko Ono, Hisao Yamada · 1988
A readability formula is obtained that can be used by computer programs for style checking of Japanese texts and need not syntactic or semantic information. The formula is derived as a linear combination of the surface characteristics of the text that are related to its readability: (1) the average number of characters per sentence, (2) for each type of characters (Roman alphabets, kanzis, hiraganas, katakanas), relative frequencies of runs (maximal strings) that consists only of that type of characters, (3) the average number of characters per each type of runs, and (4) tooten (comma) to kuten (period) ratio.To find the proper weighting, principal component analysis (PCA) was applied to these characteristics taken from 77 sample texts.We have found a component which is related to the readability. Its scores match to the empirical knowledges of reading ease. We have also obtained experimental confirmation that the component is an adequate measure for stylistic ease of reading, by the cloze procedure and by the examination on the average time taken to fill out one blank of the cloze texts.