Query Performance Analysis of NoSQL and Big Data

Ashis Kumar Samanta, Bidut Biman Sarkar, Nabendu Chaki · 2018

In the present era, Data Science is so often used for prognostic analysis instead of analyzing non-volatile data. Data has become important for research, industrial commerce and technology. Many of the repositories hold unstructured, voluminous data of different varieties generated at a high velocity. These include structured data (SD), semi-structured data (SSD) and unstructured data (USD) and collectively often referred as Big Data. The varieties of Big Data include text, audio, video, graphics and multimedia data. One of the important relevant issue is storage and retrieval of Big Data at a reasonably fast pace. NoSQL is a schema less approach that claims to handle this huge volume and variety of data efficiently. NoSQL can provide both vertical and horizontal scaling whereas conventional relational data models supports only vertical scaling. MongoDB and Cassandra are examples of NoSQL data models. The strength of our study is on hands-on implementation with multiple Big Data products used in the industry for last several years. In this empirical study, we present performance analysis of a few standard queries on SQL Server 2012, Cassandra and MongoDB data models. We make a comparative analysis of the access time on these platforms models by calculating the correlation coefficient between the number of row and access time with the increasing number of rows exceeding millions of data rows. The detailed study and interpretation of the experimental findings done in this paper would eventually contribute towards designing a general-purpose computation engine for Big data processing.

Read the paper · More papers on PaperTik