A K-NN based Data Reduction Technique in String Space via Space Separation

Rajat Kishor Varshney, Sanjay Pratap Singh Chauhan, Vishnu Sharma · 2021 3rd International Conference on Advances in Computing, Communication Control and Networking (ICAC3N) · 2021

The two most frequent representations in Pattern Recognition are statistical coding which characterize objects as vectors and structural depictions encoding components as symbolic high-level data structures like strings, trees or graphs. The two most frequent representations in Pattern Recognition are statistical codifications and structural representations. Although the overwhelming majority of classifiers can handle statistical spaces, only a few techniques can handle structural representations. The kNN Classifier, which is a neural network classifier, is one of the few techniques for both statistical and structural areas. Because it is based on the calculation of differences between all samples in the set, this approach is very versatile, but also inefficient. One of the solutions to this problem is to utilise prototype creation. These methods produce a condensed version of the original dataset by transforming and aggregating the detain the original collection. While these generating methods are very reliant on the data format, they are not well suited for structural data production. The generation- based reduction method was utilized in this case. In this article, we show how string data may be reduced to a more manageable size using homogenous clusters. This method divides space into class- homogeneous clusters, after which the median of each group is a representative prototype. The first step to address this problem is to find the median element in a string collection. According to our comprehensive test results, our method exceeds competitors in both statistical and string-based fields. We were able to demonstrate the effectiveness of our strategy using our technology to exhibit competitive differences in the categorization rate and the data reduction.

Read the paper · More papers on PaperTik