Pattern Matching Machines for Japanese Texts
Takeshi Shinohara, Setsuo Arikawa · Kyushu University Institutional Repository (QIR) (Kyushu University) · 1986
Texts in Japanese use many characters, Japanese alphabet kana and Chinese letter ha.nji, unlike texts in European languages. For that reason, Japanese characters are represented by 2-byte code in most computer systems. In many cases, the usual I-byte characters are used together '.vith 2-byte characters. In this paper, we discuss pattern matching algorithms for Japanese texts, in which I-byte characters and 2-byte characters are mixed. We have already succeeded to realize run-time efficient pattern matching machines for texts of I-byte characters by dividing character codes. We show that the method of dividing character codes is also applicable to pattern matching machines for Japanese texts. In text processing it is essential to quickly locate some or all occurrences of keyWords in text. Such techniques are usually called pattern matching(l]. The most important techniques are the ones by Knuth-Morris-Pratt[IO], Boyer-