A ProbabiUstic Approach to Japanese Lexical Analysis
Virginia Teller, Eleanor Olds Batehelder · 1993
We report on a project to develop a stochastic lexical analyzer for Japanese and to compare the accuracy of this approach with the results obtained using conventional rule-based methods. In contrast with standard, knowledge intensive methods, the stochastic approach to lexical analysis uses statistical techniques that are based on probabilistic models. This approach has not previously been applied to unrestricted Japanese text and promises to yield insights into word formation and other morphological processes in Japanese. An experiment designed to assess the accuracy of a simple statistical technique for segmenting hiragana strings showed that this method was able to perform the task with a relatively low rate of error. A debate is being waged in the field of machine translation about the degree to which rationalist and empiricist approaches to linguistic knowledge should be used in MT systems. While most participants in the debate seem to agree that both methods are useful, albeit for different tasks, few have compared the limits of knowledge based and statistics based techniques in the various stages of translation.