Method of Classification through Normal Distribution Approximation Using Estimating the Adjacent and Multidimensional Scaling
Nobuyuki Kobayashi, Takayoshi Mihara, Hiromitsu Shiina · 2016
A Two types of classification methods are applied performing classification using machine learning: (1) those for which there is a presumption of data distribution using kernel functions, such as in support vector machine (SVM) and (2) those for which data distribution is not presumed, such as the k-NN. With such methods, it is easy to obtain data whose values are close for the class as a whole. In addition, it is assumed that there is a high probability that data with close values apparently belong to the same class. In contrast, while it may be easy to obtain approximate values of data for each class, there exists a relatively high probability of approximate values manifesting as data of the same class. Consequently, there is a high probability of the same class of data appearing around the position of parameters of the data manifested in feature space where data parameters are obtained. This study proposes machine learning algorithms that approximates this using a density function of normal distribution. In addition, if small amounts of data of a different class exist where data of the same class is being collected, the impact of those classes will be significant. This study propose two types of algorithms. One is a method of calculating the influence of the proximity of the training data for the entire feature space calculates. The other is a method of calculating the effect on the entire feature space for each training data. Both methods utilize two parameters as the influence of each training data and the range to be used as neighborhood data. Two parameters determine the quasi-optimal solution by the steepest descent method. Furthermore, in order to reduce the density of the influence of the training data, we propose improved method that relocates the training data from the distance between the training data by multidimensional scaling as preprocessing.