Ranking Mutual Information Dependencies in a Summary-based Approximate Analytics Framework

Dominik Ślȩzak, Janusz Borkowski, Agnieszka Chądzyńska-Krasowska · 2018

We continue our research on utilizing histogram¬based data summaries in approximate derivation of mutual information scores in large relational data sets. Our methodology of creating, storing and using summaries has been designed for the purpose of developing an approximate database engine that is currently deployed commercially in the area of cyber- security data analytics. However, a similar idea of approximate data processing operations can be considered also in other fields, including machine learning whereby heuristic calculations are a component of many methods. In this paper, we focus on investigation of one possible source of inaccuracy of our previ¬ously proposed approach to approximating mutual information - that is, neglecting a kind of column domain drift during distributed summary-based computations. We illustrate it using an artificially created benchmark data set and we discuss how to cope this particular challenge in the future.

Read the paper · More papers on PaperTik