Building a Python library for anonymizing sensitive data

Esmeralda Madrazo Quintana · UCrea (University of Cantabria) · 2023

Technologies that handle large amounts of data have experienced rapid growth in recent years, thanks mainly to the easy availability of large volumes of data (big data). Problems arise when trying to maintain the balance between privacy and preserving as much information as possible. The dilemma of privacy preservation is further intensified when handling databases containing, for example, clinical patient data. The objective of this master’s thesis is to address privacy issues in data science by exploring and implementing the most common anonymization techniques. More specifically, we intend to implement a Python library with the most popular anonymization models, more specifically k-anonymity, l-diversity and t-closeness, as well as offer some performance analysis techniques for its optimal implementation.

Read the paper · More papers on PaperTik