LC-SID: Developing a Local LLM-based Chain-of-thought Framework for Enhanced Sensitive Information Detection
Aohai Zhang, Zhe Sun, Lihua Yin, Shaoying Li, Yutong Liang, Chao Li, Jing Du, Yanwei Sun · 2024
The detection of sensitive personal information has long been a significant area of research. However, existing methods have inherent limitations. For instance, rule-based matching requires substantial human involvement in the design process and is not well-suited for unstructured documents. Additionally, algorithms based on neural networks, such as LSTM, are limited by the type and size of the training dataset, which restricts their versatility and applicability across diverse domains. Large language models (LLMs) demonstrate robust language comprehension capabilities, making them promising candidates for detecting sensitive personal information. However, directly using mature commercial models on the network may risk leaking sensitive information. This paper proposes an LLM-based framework for identifying sensitive information, which can be deployed locally using open-source LLMs. Our proposed LCSID framework divides the recognition of sensitive information into three subtasks, utilizing a chain-of-thought (CoT) approach. To optimize resource consumption and scalability, each subtask can be handled by its own model. Ultimately, the detection rate for personal sensitive information was $91.9 \%$.