Environment for Datasets Processing and Visualization Using SciDB
Rui Wu, Sergiu M. Dascalu, J. S. Harris · 2015
When a scientist analyzes certain datasets, there are three common steps—import, visualize, and process data. There are some prevalent tools to visualize and process data, such as Matlab, Powersim, and Stella. However, these tools cannot handle big data. For data management, scientists are pursuing a new generation of Database Management Systems (DBMS) to replace traditional Relational Database Management Systems (RDBMS), because RDBMS is good at data management, but does not perform well on raw data and time series data. Furthermore, most of the new tools, such as Hadoop, cannot fulfil scientists’ needs of data management. This paper introduces a web-based application created by us to import, visualize, and process big data. The system offers two methods for users to import data—obtain data from foreign repositories and upload users’ files. We used SciDB for data management, because SciDB is designed for big data and scientific use. We used D3.js and Dygraphs libraries for data visualization, which enable users visualize millions of points without experiencing lag.