Identifying Near-Duplicate Pairs in Exams: Similarity Detection Using Kshingling, Minhashing, and LSH
Nabila El Rhezzali, Imane Hilal, Hnida Meriem · 2023
In the field of data mining, identifying near-duplicate documents is a key challenge. This task, crucial for detecting cheating in online exams, can be approached using techniques from Natural Language Processing: Shingling and Minhashing. Furthermore, using the Locality Similarity Hashing (LSH) approach refines the process by focusing on high-probability document comparisons. This paper introduces a novel method that combines K-shingling, Minhashing, LSH, to uncover near-duplicate pairs in online exams. Experimental results highlight the success of this integrated approach in enhancing detection accuracy, ultimately strengthening the capability to address academic dishonesty.