Efficient semantically equal join on strings
Juggapong Natwichai, Xingzhi Sun, Maria E. Orłowska · 2007
Data integration issues are again the focus of many research groups who now benefit from the accumulated experience of the last two decades. As in the past, data warehousing was the main driver for data unification, and now e-research (e-science) applications, spanning multiple sides of wide computer networks, add to this call to investigate the automation of data integration. The data fusion process only begins with schema integration and must be followed by detailed data instances resemblance in order to reach data representation unification. The data mismatch can be observed due to many reasons such as typing errors, different abbreviation conventions, different standards for data representation and coding, etc. Hence data cleaning and redundancy removal are unavoidable procedures in overall data integration processing. In this paper we address the limits of the data reconciliation automation process in cases where the compared data is semantically equivalent but the data native representation of the values of given attributes is different. We assume that a semantic