Use of Machine Learning for the Unknown Values in Database Transformation Processes
Michal Kvet, Roman Čerešňák, Veronika Šalgová · 2021
Relational databases have fulfilled the role of a primary data repository for decades. With the development of modern technologies, the type of databases has come to the forefront which - unlike relational databases - have an unconventional structure which allows greater flexibility in data storage. Since the design and implementation of non-relational databases, which dates to the end of the 20th century, various types of non-relational databases have been proposed - the most popular non-relational databases include document, graphic, column and key-value databases. Although the flexibility of non-relational databases is an excellent feature in many respects, the freedom of the data structure and the freedom of defining data type cause considerable problems in number of database aspects. One of the situations in which the freedom of the data structure causes problems is the transformation of non-relational databases (the most common document database MongoDB or DynamoDB) to relational or other type of database. Undefined values often cause a termination of a running process and/or higher error rate in the transformation process. For these reasons, it is necessary to solve problems with undefined values either before or during the process of transformation and thus avoid application errors. To this end, we have created a module for solution of the problem of undefined values in the transformation of a non-relational database MongoDB to a relational database Oracle. The main purpose of this article is not to eliminate errors when starting the process, but to design and implement a module that can apply a dynamic solution based on input data and the selection of the target database when unknown values occur in the non-relational database MongoDB.