Detecting and Protecting Personally Identifiable Information through Machine Learning Techniques

Carlos Jorge Augusto Pereira da Silva · Open Repository of the University of Porto (University of Porto) · 2020

This dissertation is a study on the automatic detection of Personal Identifiable Information (PII) in emails.The definition of PII has evolved following the technological developments and the misuse of PII.PII storage has become easier, because of the diminishing costs and accelerated digitization of business.Nevertheless, the security aspects of data storage have been neglected resulting in an increasingly common availability of PII on the internet inadvertently taken from companies.The digitization of business also contributed to the realization that there was value to be extracted from clients PII.Creating user profiles, adapting campaigns, and other techniques added value to their business, making companies willing to storage as much information as possible about their users, even if they had no use for it at the moment.For companies that had a software platform, there was the realization that the more data the customer inserted in the system the harder it would be for the customer to leave for a competitor.So, they made it difficult, or even impossible, to extract the information and move to another system, increasing customer retention rates.Each European country had their own legislation on data protection, making it more difficult for companies to run simultaneously in multiple European countries.It is in this context that in 2018 the General Data Protection Regulation (GDPR) comes into force in the European Union (EU), looking into address these issues.Creating a single legislation across all European Union countries, promoting the data security best practices to protect consumers, and regulating the circulation of PII across the EU countries and companies in them.An element that is common to most companies is the significant quantity of information that is stored and in circulation in emails.Including personal data of clients, employees, and collaborators.To follow the GDPR it is required an automated solution that allows the detection of PII at the high rate of email transactions.In this dissertation we investigate how different models and Natural Language Processing techniques help to achieve that goal.The focus of these investigations is the design of a microservice that receives text content and detects PII named entities.For that, we trained machine learning models for PII detection and for segmenting emails into parts.We found that the state-of-the-art techniques are too expensive to run in production environments but that a good alternative could be achieved even if not achieving state-of-the-art results.

Read the paper · More papers on PaperTik