Graph Analytics on Massive Collections of Small Graphs

Dritan Bleco, Yannis Kotidis · 2014

Emerging applications face the need to store and query data that are naturally depicted as graphs. Building a Business Intelligence (BI) solution for graph data is a formidable task. Relational databases are frequently criticized for being unsuitable for managing graph data. Graph databases are gaining popularity but they have not yet reached the same maturity level with relational systems. In this pa-per we identify a large spectrum of applications that generate graph data with specific characteristics that make them candidate for be-ing stored in a relational system. We describe a novel framework where data and queries are both treated as abstract graph structures that can be decomposed into simpler structural elements. We com-plement this abstract framework with a description of a system that utilizes three different means of expediting user queries: (i) a flat description of the graph records using a column-oriented storage model, (ii) use of bitmap columns for enabling fast access to parts of these graph records and (iii) a novel framework for selecting and materializing graph views that significantly expedite retrieval of records in response to a graph query. To the best of our knowl-edge we are the first to report results using datasets consisting of hundreds of millions of graphs with billions nodes, edges and mea-sure values using a single database server running of a commodity node. Our results demonstrate that our platform is orders of mag-nitude faster than alternative systems that natively handle graph data and a straightforward relational implementation. Moreover, our materialization techniques (that account for about 10 % of extra disk space) are able to reduce the query execution times further, by up to 94 % compared to an evaluation plan that is oblivious to the existing materialized graphs views in the database. 1.

Read the paper · More papers on PaperTik