Implementation of genetic network programming and knapsack problem for record clustering on distributed database

Wirarama Wedashwara, Shingo Mabu, Masanao Obayashi, Takashi Kuremoto · 2014

This research involves implementation of genetic network programming (GNP) and knapsack problem (KP) to solve record clustering on distributed databases. The objective is to distribute big data to certain sites with the limited amount of capacities by considering the similarity of distributed data in each site. GNP is used to extract rules from big data by considering characteristics (value ranges) of each attribute in a dataset. KP is used to distribute rules to each site by considering similarity (value) and data amount (weight) related to each rule to match the site capacities.

Read the paper · More papers on PaperTik