Method for creating privacy-preserving information using Probabilistic Latent Semantic Analysis

Yuki Sugawara, Eiichi Sakurai, Yoichi Motomura, Yukihiko Okada, Kai Tanabe, Akiko Tsukao, Shinya Kuno · 2023

In recent years, rapid digitization and technological advances have enabled governments and companies to utilize various types of personal information, including environment, medical, and lifestyle. Privacy protection is an important issue in the utilization of personal information, and anonymization methods such as k-anonymization and 1-diversity have been proposed in the past. However, to utilize data effectively, it is also important not to lose the usefulness of the information. Privacy protection and preserving the usefulness of information are in a trade-off relationship, and balancing these two aspects has been a difficult problem. Recently, microaggregation using clustering algorithms has been attracting attention as an anonymization method that ensures the usefulness of the data while protecting privacy. In this study, we proposed a new microaggregation method based on iterative processing using probabilistic latent semantic analysis and evaluated the proposed method based on the use case. The use case is a classification of social isolation and an analysis of the relationship with physical frailty risk using the classified groups. Compared to microaggregation using quasi-identifiers and microaggregation based on a conventional clustering algorithm, the proposed method was able to find social isolation characteristics obtained in the raw data with higher accuracy. In addition, in a more detailed analysis focusing on a population with certain characteristics, the proposed method was able to replicate the analysis results obtained using the raw data and to make appropriate assessments of the health risks.

Read the paper · More papers on PaperTik