A cost sensitive classifier for Big Data
Akash Neelesh Haldankar, Kiran Bhowmick · 2016
Data Mining techniques have been used to detect fraud related to several domains like risk identification. An assumption about the data is that it is always balanced, this is far from true. It doesn't represent the reality. In this paper we develop a cost sensitive classifier to detect Risk using the Statlog (German Credit Data) data set. This study shows how application of proper feature selection followed by using a unique combination of ensemble & thresholding helps to reduce the overall cost. We also see the effects of this classifier on unstructured data as well as streaming data.