Detecting Translingual Plagiarism And The Backlash Against Translation Plagiarists
Rui Sousa‐Silva · Portuguese National Funding Agency for Science, Research and Technology (RCAAP Project by FCT) · 2014
Plagiarism detection methods have improved signi cantly over the last decades, and as a result of the advanced research conducted by computa- tional and mostly forensic linguists, simple and sophisticated textual borrowing strategies can now be identi ed more easily. In particular, simple text compari- son algorithms developed by computational linguists allow literal, word-for-word plagiarism (i.e. where identical strings of text are reused across di erent docu- ments) to be easily detected (semi-)automatically (e.g. Turnitin or SafeAssign), although these methods tend to perform less well when the borrowing is obfus- cated by introducing edits to the original text. In this case, more sophisticated linguistic techniques, such as an analysis of lexical overlap (Johnson, 1997), are required to detect the borrowing. However, these have limited applicability in cases of ‘translingual’ plagiarism, where a text is translated and borrowed with- out acknowledgment from an original in another language. Considering that (a) traditionally non-professional translation (e.g. literal or free machine trans- lation) is the method used to plagiarise; (b) the plagiarist usually edits the text for grammar and syntax, especially when machine-translated; and (c) lexical items are those that tend to be translated more correctly, and carried over to the derivative text, this paper proposes a method for ‘translingual’ plagiarism detec- tion that is grounded on translation and interlanguage theories (Selinker, 1972; Bassnett and Lefevere, 1998), as well as on the principle of ‘linguistic unique- ness’ (Coulthard, 2004). Empirical evidence from the CorRUPT corpus (Corpus of Reused and Plagiarised Texts), a corpus of real academic and non-academic texts that were investigated and accused of plagiarising originals in other languages, is used to illustrate the applicability of the methodology proposed for ‘translingual’ plagiarism detection. Finally, applications of the method as an investigative tool in forensic contexts are discussed.