Efficient Outlier Detection for High Dimensional Data using Improved Monarch Butterfly Optimization and Mutual Nearest Neighbors Algorithm: IMBO-MNN
M. Rao Batchanaboyina, Nagaraju Devarakonda · International journal of intelligent engineering and systems · 2020
The hybrid Improved monarch butterfly optimization-mutual nearest neighbor (IMBO-MNN) is proposed for outlier detection in high dimensional data.It is a challenge to detect outliers in high dimensional information.The external behavior of the data points cannot be detected in high-dimensional data except in the locally relevant data sub-sets.Subsets of dimensions are called subspaces, and with an increase in data dimension, the number of those subspaces grows exponentially.In another subspace an information point that is an outlier can appear ordinary.It's essential to assess its outlier behavior according to the amount of subspaces in which it appears as an outermost part to characterize an outlier.Data is scarce in high-dimensional space and the concept of closeness does not preserve meaning.In fact, the sparsity of the high-dimensional data means that every point is nearly equal from the point of view of closeness-based finishes.As a result, for higher dimensional information finding is more complicated and non-obviously significant outliers.An enhanced MBO (IMBO) algorithm is offered for enhanced search precision and run time efficiency by a fresh adaptation provider.Statistical results indicate that the elevated local optimal prevention and quick convergence rate of the improved monarch butterfly optimization (IMBO) algorithm helps to exceed the basic MBOs in outlier detection.Comparatively, IMBO produces very competitive outcomes and tends to surpass present algorithms.Optimal value k remains a task, affecting the efficiency of kNN straightforwardly.We are presenting a fresh learning algorithm under kNN in this paper to alleviate this issue called mutual nearest neighbor (MNN).The main feature of our method is that the class marks of unknown instances are defined by mutually next to one another, instead of by closest neighbor.The advantage of mutual neighbors is that in the course of the prediction process pseudo close neighbors can be identified and taken not into account.The performance of the suggested algorithm has been examined with a number of studies.For 100 data, IMBO-MNN is in 4897 milliseconds, PSO is in 5239 milliseconds, random forest is 5347 milliseconds, PNN is 5278 milliseconds and KNN is in 5166 milliseconds.For 250 data, IMBO-MNN is in 5984 milliseconds, PSO is in 6132milliseconds, random forest is 6145 milliseconds, PNN is 6124 milliseconds and KNN is in 6152 milliseconds.For 500 data, IMBO-MNN is in 6416 milliseconds, PSO is in 6636 milliseconds, random forest is 6634 milliseconds, PNN is 6719 milliseconds and KNN is in 6710 milliseconds.For 1000 data, IMBO-MNN is in 6913 milliseconds, PSO is in 7111 milliseconds, random forest is 7019 milliseconds, PNN is 7134 milliseconds and KNN is in 7162 milliseconds.The proposed IMBO-MNN performs better with minimum time taken.The findings indicate that the technique can identify outliers in high dimensional data efficiently in a decreased calculation moment.