Chain-of-Thought Enhanced Content Detection in Large Language Models

Chenjun Zou, Zhuoyan Feng, Hongwei Guan, Chuanjun Xing · 2024

This paper proposes and explores a content detection method for large language models based on the Chain of Thought (CoT) logic chain. By incrementally extracting logic chains from text or code and converting them into feature vectors, this method enables logical plagiarism detection and content verification for complex tasks. Particularly for classical algorithmic problems, such as dynamic programming, this approach effectively mitigates plagiarism attempts through variable renaming. The paper first delves into the theoretical foundation of CoT logic chains and analyzes their application mechanisms within large language models. It then provides a detailed account of the steps for feature extraction and similarity detection algorithms, including how to efficiently convert logic chains into feature vectors for comparison. To validate the effectiveness of this method, several classical algorithmic instances were selected for experimentation. Results indicate that the CoT logic chain-based content detection method offers significant advantages in detecting originality in academic papers and code. This approach not only substantially improves the detection rate of academic misconduct but also assists editors and reviewers in more accurately identifying AI-generated content, thereby holding broad applications in academic publishing, code review, and educational fields. Additionally, the paper discusses the potential application of this method in other domains, such as patent plagiarism detection and legal document similarity analysis.

Read the paper · More papers on PaperTik