Using Political Party Affiliation Data to Measure Civil Servants' Risk of Corruption
Ricardo Silva Carvalho, Rommel Novaes Carvalho, Marcelo Ladeira, Fernando Mendes Monteiro, Gilson Libório Mendes · 2014
This paper presents a case study of machine learning applied to measure the risk of corruption of civil servants using political party affiliation data. Initially, a statistical hypothesis test verified the dependency between corruption and political party affiliation. Then, we constructed datasets with standardization and three different discrimination techniques. Using Weka environment, this work shows the application and statistical evaluation of four classification algorithms to build models for predicting risk of corruption: Bayesian Networks, Support Vector Machines, Random Forest, and Artificial Neural Networks with back propagation. To evaluate the models we used data mining metrics such as precision, recall, kappa statistic and percent correct. Lastly, the case study compares the learned model with the best performance to the experts' model. The comparison not only confirms previous experts' affirmations, but also provides new assertions on the affiliation-corruptibility relation.