Clustering a translational corpus

Shih‐Wen Ke · Studies in corpus linguistics · 2012

This chapter describes the various clustering techniques and document processing methods one can use to discover information about similarities found in translational corpora. Two types of clustering techniques, namely hierarchical clustering and partitioning clustering, and their variations are discussed and applied to a sample of the TK-NHH Translatørkorpus corpus consisting of 71 translated documents on 4 different topics. The results show that these clustering techniques are capable of differentiating translations accepted by experts from those rejected, suggesting that these accepted translations share a high degree of similarity and perhaps resemble an ideal translation of the original text.

Read the paper · More papers on PaperTik