A Knowledge Management Framework for imbalanced data using Frequent Pattern Mining based on Bloom Filter

Sally M. Elghamrawy · 2016

Managing medical environments and organizations performance depend directly on the knowledge management (KM) systems. Knowledge Discovery (KD) is responsible for digging information from datasets and finding internal knowledge within organizations or external sources. Data mining (DM) is the core of KD process. Although recent mining techniques have proven their accuracy in discovering the knowledge from balanced data, where the class distribution is balanced, the problem of discovering knowledge from unbalanced data is still a challenge that needs to be addressed. A Clustered Knowledge Management Framework (CKMD) is presented in this paper, for enhancing the performance of KD from unbalanced data. A Simple Hybrid Sampling Approach (SHSA) is proposed to reduce the adverse impacts of imbalanced data. Mining frequent pattern process plays an important role in KD process. Moreover, a Frequent Pattern Mining algorithm based on Bloom Filter (FPMBF) is proposed to discover items that frequently co-occur in the data using the bloom filter, that requires a single scan of the data, which leads to less time consuming in discovering knowledge for imbalanced data. Finally, the performance of the proposed methods is evaluated using real datasets and comparative experiments.

Read the paper · More papers on PaperTik