Lexical similarity using fuzzy Euclidean distance
Heba Ayeldeen, Aboul Ella Hassanien, Aly A. Fahmy · 2014
Knowledge exaction and text representation are considered as the main concepts concerning organizations nowadays. The estimation of the semantic similarity between words provides a valuable method to enable the understanding of texts. In the field of biomedical domains, using Ontologies have been very effective due to their scalability and efficiency. In this paper, we aim to cluster and classify medical thesis data to better discover the commonalities between theses data and hence, improve the accuracy of the similarity estimation which in return improves the scientific research sector. Experimental evaluations using 4,878 theses data set in the medical sector at Cairo University indicate that the proposed approach yields results that correlate more closely with human assessments than other by using the standard ontology (MeSH). Two different algorithms were used; the first is Lexical similarity and then applying K-means clustering and the second is fuzzy Euclidean distance clustering algorithm after using MeSH ontology on medical theses data for better categorization of the keywords within the data.