PrivGraph: Modeling Personal Privacy Information Using Knowledge Graph
Jinhui Zuo, Seokwon Lee · 2025
Prior research has demonstrated that Artificial Intelligence (AI) models can memorize specific training data, which can then be extracted through targeted attacks. Furthermore, AI's inherent ability to integrate information allows for the aggregation of disparate data points from diverse sources, creating significant privacy vulnerabilities. Conversely, existing Privacy-Enhancing Technologies (PETs) treat the entire dataset or individual Personally Identifiable Information (PII) as uniformly sensitive, neglecting the inherent relationships and interactions between data elements. To address this gap, we propose that data should be classified as sensitive only when a sufficient quantity of such information pertaining to an individual is present within the data. To this end, this study employs knowledge graphs to model private information, creating a hierarchical structured representation, termed PrivGraph, that reveals sensitive data aggregation within training datasets. Leveraging PrivGraph, we introduce the Sensitive Level Factor (SLF) to quantify the degree to which an individual's private information is embedded in the data. Furthermore, we propose PrivGraph-based knowledge probing to facilitate posttraining privacy assessments. Our experiments demonstrated that PrivGraph achieves comparable performance to existing models in private information detection, while also effectively modeling individual private information, even with lengthy texts and data from diverse sources. At the end of the paper, we discuss how PrivGraph can be incorporated throughout the AI engineering lifecycle to provide traceable as well as comprehensive privacy protection. Our code is available at https://github.com/pphui8/PrivGraph.