Scalable data perturbation for privacy preserving large scale data analytics and machine learning
MAHAWAGA ARACHCHIGE, Pathum Chamikara · Figshare · 2021
Advancements in the Internet and related technologies such as the Internet of Things (IoT) are transforming the major industries such as healthcare, banking, agriculture, energy, and transportation. As cyberspace permeates physical and human spheres, a vast amount of data (e.g. big data, data streams) becomes available for analysis. With incremental and fast data generation, data analytics and machine learning (ML) are producing remarkable accuracy in generating essential insights when presented with massive amounts of data. For example, an approach that has gained more attention is deep learning, which produces exceptional performance when trained with large amounts of data. However, industries such as healthcare, banking, agriculture, energy, and transportation closely work with sensitive person-specific private data. There is a growing concern that potentially sensitive data may become public if the collected data are not appropriately sanitized before they are released for investigation. Additionally, organizations often want their insights to be restricted within the organizational boundaries and maintain the trustworthiness of the data and knowledge. Data analytics and ML must be equipped with privacy-preservation scenarios so that user privacy is intact while the data analytic and ML knowledge is disseminated. Although there are more than a few privacy preservation approaches, the high dimensionality of big data and data streams makes privacy preservation challenging, making the existing privacy preservation approaches obsolete. The central issue of the existing approaches is their incapability to maintain the right balance between privacy, utility, and efficiency when introduced to high dimensional data such as big data and data streams. Some of the efficient approaches provide good privacy but fail to provide good utility, whereas other efficient approaches offer good utility but fail to provide good privacy. In this thesis, we looked at the problem of preserving the privacy of large scale data analytics and machine learning under four separate categories: C1. Privacy preservation of big data, C2. Privacy preservation of data streams, C3. Privacy preservation of machine learning, C4. Implementing privacy in real-world settings. Under C1, this thesis explores the complexity in the privacy preservation of big data sharing (the static settings) for data analytics. In this case, the whole dataset is assumed to be already available for processing. The main challenge of C1 is the development of efficient privacy preservation approaches for high dimensional data while maintaining privacy and utility. Under C1, an efficient big data perturbation approach named PABIDOT is introduced. Next, the concepts used in PABIDOT are extended to generate another privacy preservation approach named DISTPAB, which works in distributed machine learning settings. Under C2, this thesis investigates the complexities in data stream privacy preservation (the dynamic setting). Under this category, the main challenge is the maintenance of efficiency, privacy, and utility of a data stream with a continuous incremental flow. This thesis proposes two privacy preservation approaches named PR2oCAl and SEAL under C2. PR2oCAl provides an efficient privacy preservation approach for data streams while maintaining a high level of utility as the data grows. However, C2 needed further investigation as PR2oCAl introduces a low efficiency when the number of attributes of the dataset increases exponentially (which can occur when the number of IoT sensors increases exponentially). In order to successfully address this issue, this thesis introduces SEAL, which provides linear complexity for both the number of attributes and the number of instances. Under C3, this thesis looks at privacy-preserving machine learning in large scale distributed settings. The main challenge addressed in C3 is the privacy of individuals in a distributed machine learning environment. This thesis proposes two privacy-preserving approaches named LATENT and PEEP. LATENT is a differentially private deep learning approach that presents a distributed approach to overcome the issues in server-centric privacy-preserving deep learning approaches. PEEP is a privacy-preserving face recognition approach that proposes a perturbation mechanism on the face data to avoid third-party servers from gaining access to sensitive biometric data. Under C4, this thesis introduces two comprehensive frameworks named PriModChain and PPaaS for two real-world scenarios. PriModChain aims to improve the privacy, security, immutability, and traceability of machine learning (ML) models in IIoT systems. Compared to existing approaches, PriModChian provides a feasible mechanism to share the ML knowledge between the branches of an organization with privacy and trustworthiness. PPaaS is introduced to tailor privacy preservation to stakeholder needs by reducing the complexity of choosing the best data perturbation algorithm suitable for a particular scenario (e.g. data classification). All the proposed approaches have been tested with real-world data and compared against existing approaches to show the competitive advantages.