ROBiT - A Binary Optimization Anti-Plagiarism Method

Roberta Robert, Bruno Castro da Silva, Jeferson Campos Nobre · 2025

The widespread availability of Large Language Models (LLMs) has significantly lowered the barrier to committing code plagiarism. However, most existing anti-plagiarism tools remain vulnerable to modern evasion strategies, including syntactic transformations and generative code rewriting. Prior work shows that such transformations can effectively bypass clone detectors that rely on syntactic or semantic representations. While binary optimization is a known technique in malware obfuscation, its potential for plagiarism detection has been largely overlooked. We introduce a hybrid detection method that combines source-level syntactic analysis with binary-level comparison, leveraging both standard compilation outputs and binaries generated with optimization flags. These optimizations act as a reverse filter, eliminating syntactic manipulations added to code artifacts and revealing structural similarities with the original binary. Our empirical evaluation confirms that optimized binaries exhibit patterns that correlate strongly with their original source code. The proposed method demonstrates high effectiveness in detecting plagiarism, even when the source code has undergone aggressive syntactic transformations. This technique serves as a robust and complementary extension to existing syntactic anti-plagiarism systems, offering deeper insight into semantic and structural code similarity.

Read the paper · More papers on PaperTik