An infrastructure for scalable parallel multidimensional analysis

Sanjay Goil, Alok Choudhary · 2003

Multidimensional analysis in online analytical processing (OLAP), and scientific and statistical databases (SSDB) use operations requiring summary information on multidimensional data sets. Most common are aggregate operations along one or more dimensions of numerical data values and/or on hierarchies defined on them. Simultaneous calculation of multidimensional aggregates are provided by the Data Cube operator. This is computed only partially if the number of dimensions is large. Queries may either be answered from a materialized cube or calculated on the fly. The multidimensionality of the underlying problem can be represented both in relational and multidimensional databases, the latter being a better fit when query performance is the criteria for judgement. Relational databases are scalable in size for OLAP and multidimensional analysis and efforts are on to make their performance acceptable. On the other hand multidimensional databases provide good performance for such queries, although they are not very scalable. We address scalability in multidimensional systems for analysis in SSDB and OLAP applications. We describe our system PARSIMONY-Parallel and Scalable Infrastructure for Multidimensional Online analytical processing. Sparsity of data sets is handled by using chunks to store data as a sparse set using a bit encoded sparse structure. Chunks provide a multidimensional index structure for efficient dimension oriented data accesses. Operations within and between chunks are a combination of relational and multidimensional operations depending on whether the chunk is sparse or dense. Performance results for high dimensional data sets on a distributed memory parallel machine (IBM SP-2) show good speedup and scalability.

Read the paper · More papers on PaperTik