New similarity functions

Hossein Yazdani, Daniel Ortíz-Arroyo, Halina Kwaśnicka · 2016

In data science, there are some parameters that affect the accuracy of selected algorithms, regardless of their type. Type of data objects, membership assignments, and distance or similarity functions are the most important parameters that provide or not a proper environment for learning algorithms. The paper evaluates similarity functions as fundamental keys for membership assignments. The issues on conventional similarity functions are discussed in this paper. The paper introduces Weighted Feature Distance (WFD), and Prioritized Weighted Feature Distance (PWFD) to cover diversity in feature spaces. Most of the conventional distance functions compare data objects on vector space where any dominant feature may massively skew the final results. WFD functions perform better in supervised and unsupervised methods by comparing data objects on their feature spaces in addition to covering similarity on vector space. Prioritized Weighted Feature Distance (PWFD) works as same as WFD with ability to give priorities to desirable features. The accuracy of proposed functions are compared with other similarity functions on some data sets. Promising results show that the proposed functions work better than the other methods presented in this literature.

Read the paper · More papers on PaperTik