An automatic indexing for thai text retrieval

Chuleerat Jaruskulchai, Gideon Frieder · Medical Entomology and Zoology · 1998

The Thai writing system has no punctuation to indicate words or sentences. This makes difficult both in human and machine learning to detect the word boundaries. Therefore, text segmentation is a prerequisite for any machine studies including information retrieval. Generally, previous works can be classified into purely dictionary, morphological rules, and statistical approaches. Most of the segmentation research had been done under machine translation; however, the current success is only in the word-by-word translation. Furthermore, the major task in the information retrieval which is identifying keywords, is done by experts. This research establishes the tasks of the information retrieval and test data sets for both information retrieval and segmentation algorithms. The indexing techniques, n-gram, probabilistic word-based, and morphological rules approaches for Thai text were investigated. This research showed that the most prominent approach, n-gram, did not contribute to Thai text due to the index size. The basic morphologies and simple term weights achieved a more desirable retrieval effectiveness. The most striking term weighting system was the augmented term frequencies. It re-weighted the normal term frequencies to lie between 0.5 and 1. The results of the morphological rules, on the other hand, were achieved only in the phoneme level and still miss word-level semantics. A new technique in acquiring lexicons, BayNet, was presented. The motivation of the BayNet resembles the human learning process. In this new technique, the basic morphologies were used at the first stage, then statistics were obtained for combining the basic monosyllables. The well-known Minimum Description Length (MDL) principle of Bayesian Networks was exploited in acquiring the lexicon structures. Results of this thesis provided insight for the Thai information retrieval systems. The results had shown that perfect segmentation is not a mandatory process. Thai writing system posed a unique problem in the area of information retrieval due to the characteristics of the word formations and word usage. Although the new technique mentioned above did not provide dramatically greater results than dictionary based approach, it provided extensive ideas in obtaining the new lexicons without using a dictionary as a table lookup or word patterns for learning.

Read the paper · More papers on PaperTik