Data safekeeping against knowledge discovery

Zbigniew W. Raś, Seunghyun Im · 2006

This dissertation examines an important issue of data mining: how to provide meaningful knowledge without compromising data confidentiality. In other words, information systems should provide knowledge extracted from their data which can be used to identify underlying trends and patterns, but the knowledge should not be used to reveal confidential data. In conventional database systems, data security is to be maintained by not revealing them to unauthorized users. However, hiding secret data is not sufficient in knowledge discovery systems (KDS) due to Chase [Ras and Dardzinska, 2005b] [Im and Ras, 2005]. Chase is a widely used data mining technique that helps predict missing values in information systems using knowledge provided through pattern or rule extraction. For example, if an attribute is incomplete in an information system we can use Chase to approximate the missing values to make it more complete. It is also used to answer user queries containing non-local attributes [Ras and Joshi, 1997]. If attributes in queries are locally unknown, we search for their definitions from KDS and use the results to replace the non-local part of the query. The problem of Chase with respect to data confidentiality is that it has the ability to reveal hidden data. Sensitive data may be hidden from an information system, for example, by replacing them with null values or encrypting them with cryptographic technologies. However, any user in the KDS who has access to knowledge is able to interpret hidden data as incomplete or missing elements and reconstruct them with Chase. For example, in a standalone information system with a partially confidential attribute, knowledge extracted from non-confidential part can be used to reconstruct hidden (confidential) values [Im et al., 2005b]. When a system is distributed with autonomous sites, knowledge extracted from both local or remote information systems [Im et al., 2005b] may reveal sensitive data to be protected. Clearly, mechanisms that would protect against these vulnerabilities have to be implemented in order to build a security-aware KDS. This paper presents protection algorithms for minimizing disclosure of confidential data and describes how to build a secure KDS which best meet specific additional requirements.

Read the paper · More papers on PaperTik