Vertical query-join benchmark in a cloud database environment

Jens Köhler, Thomas Specht · 2014

Nowadays, enterprises across all branches and sectors face a new hype regarding “Big Data”. Thus, new requirements in the context of Business Intelligence emerge. Big Data demands to process vast amounts of unstructured data from social networks, sensor data, etc. in near real-time. In order to tackle these challenges, current research works aim to develop new ways of data storage and analysis from a database point of view. This is the advent of so-called “In-Memory” databases (e.g. SAP HANA) that hold entire data volumes in their fast RAM memory and use hard disks only for logging or archiving purposes. Another promising technology with respect to this topic is "Cloud Computing". Storing and analyzing vast amounts of heterogeneous data require appropriate underlying hardware infrastructures. Obtaining such hardware capabilities form external cloud providers is an auspicious way to avoid expensive investments in new hardware. However, using external hardware resources from the public cloud always means that crucial data has to leave the internal enterprise network and enterprises have to trust external providers. Bringing "Big Data" into the cloud, our approach follows the principle of vertically distributed database tables. The main idea is to divide crucial database data and distribute it across different (public and private) cloud providers. Thus, every provider only gets a small part of the data. These individual small parts are worthless without the other parts and enable enterprises to meet their compliance rules concerning data security and protection. So Cloud Computing becomes an interesting alternative to store vast amounts of data. This work evaluates our approach from a performance point of view and presents the corresponding query times with and without vertically partitioned data.

Read the paper · More papers on PaperTik