Large-scale multilingual document clustering(本文)

和明 岸田 · Institutional Repositories DataBase (IRDB) · 2013

It is often necessary to categorize automatically multilingual document sets, in which documents written in a variety of languages are included, into topically homogeneous subsets, such as when applying an automatic summarization system for multilingual news articles.However, there have been few studies on multilingual document clustering (MLDC) to date.In particular, it is not known whether clustering techniques are effective in large-scale multilingual document sets.When the target multilingual document set is enough small, it is easy to partition it automatically by using machine translation (MT) software and standard statistical package after text processing.In contrast, for other situations where large document collections have to be processed, it is necessary to explore a 'scalable' technique of MLDC because computational complexity of executing MT and DC increases exponentially, not linearly, as size of target data becomes larger.The purpose of this thesis is to develop a method of large-scale MLDC and to verify its effectiveness experimentally.The approach of this thesis for solving the large-scale MLDC problem is to combine cross-lingual information retrieval (CLIR) technique and clustering technique for large-scale document collections.For it, this thesis reviews CLIR methods and document clustering algorithms exhaustively to identify useful techniques for large-scale MLDC in terms of efficiency.As a result, the thesis adopts a combination of dictionary-based translation method in CLIR and leader-follower clustering (LFC) algorithm for implementing the MLDC system.After the reviews, results from three experiments for clarifying empirically the effectiveness of the proposed system are reported, more specifically, (a) effectiveness of scalable techniques for term (translation) disambiguation used in the dictionary-based translation, (b) effectiveness of the LFC algorithm for monolingual DC in comparison with theoretically sophisticated techniques, and (c) effectiveness of the LFC algorithm with dictionary-based translation for large-scale MLDC, by using some test collections.In experiment (a), it was observed that 'best cohesion' method for translation disambiguation works well although its computational complexity is relatively low, and in experiment (b), it was shown that the LFC algorithm for which the target file is scanned only twice can generate 'good' cluster sets, which are comparable with those obtained by the spherical k-means algorithm and the hierarchical Dirichlet process (HDP) mixture model.Finally, it was clarified in experiment (c) that the MLDC system works well for a document collection including over 13,000 news articles written in English, French, German and Italian.Through the experiments, the effectiveness and efficiency of the proposed method for MLDC were empirically confirmed.

Read the paper · More papers on PaperTik