Parallel Corpora Construction Based on Computer Technology
Ronggen Zhang · 2023
The paper processes and analyzes the parallel corpus texts of Chinese versions of the English novel Pride and Prejudice, by constructing two parallel corpora of its Chinese versions and the original text as a reference corpus. The tools used to process the data include ctbparser_0.11, PatCount, EditPad Pro, AntConc, SPSS 19, and the like. Two findings are shown as follows: first, either Chinese version of Wang or Zhang is significantly different from the original text, in that the Chinese version is filled with more adverbs, cardinal numerals, common nouns, punctuation marks, and fragment sentences; while the original text is full of more adjectives, conjunctions, pronouns, copulative verbs which are used in progressive tense and passive voice. Second, compared with Zhang’s version, Wang’s is nearer to Austen’s text, especially in the frequency of uses of adjectives, pronouns, copula verbs, and passive marker suo, i.e. by, that is, Wang’s version is more foreignized than Zhang’s. Wang’s version is more foreignized than Zhang’s. Nevertheless, either version has its merits; it can be said in such a way that Wang’s version is slightly quaint; while Zhang’s version is very vivid.