Simple and Efficient Identification of Personally Identifiable Information on a Public Website
Caitlin Brown, Charles Morisset · 2022 IEEE International Conference on Big Data (Big Data) · 2022
Personally Identifiable Information (PII) is a key concept in privacy regulation. This form of information can provide revealing information about individuals, which may be collected and used for malicious purposes, such as social engineering and identity theft [1]. Consequently, privacy preserving legislations, such as GDPR place the responsibility of appropriately handling PII onto organisations who may process large amounts of personal data as part of their day to day operations. Therefore, it is necessary to develop processes that can provide privacy assurances but also to do so with increased automation and reliability [2]. This work will focus on assessing the ability of the Natural Language Processing tool, sentiment analysis and text classification algorithms to detect PII automatically, reliably and without too much complexity. To achieve this a dataset containing web pages from Newcastle University’s website, with a focus on staff profiles was created and manually labelled to indicate which sentences contained PII. The dataset was then used to train three text classification algorithms: Multinominal Naïve Bayes, Random Forest Classifier and LSTM model in order to predict the labels of an unseen portion of the dataset. The algorithms all performed well at detecting PII, with Random Forest achieving the highest accuracy at 96% and 96% F1-Score. Nevertheless, the models all mislabelled more sentences containing PII as not containing PII, than those which did not contain PII but were labelled as doing so.