Using Fuzzy Equivalence Relations to Model Position Specificity in Sequence Kernels
Mani Manavalan · Zenodo (CERN European Organization for Nuclear Research) · 2019
One of the most fundamental issues in computational biology is the classification of biological sequences, particularly nucleotide (DNA and RNA) and amino acid (protein) sequences. Sequence statistics were widely used in the early days of computational biology. Hidden Markov Models (HMM) and Artificial Neural Networks (ANN) were used to improve these methods subsequently (ANN). SVM classifiers are essentially linear classifiers in their most basic form. SVMs differ from other linear classifiers such as perceptrons or logistic regressors in that they maximize the margin between the classes to define the classification function – a principle that is known to be optimal in terms of generalization error limitations. This paper replicates previously published sequence kernels based on the occurrence of specified patterns as well as some position-specific variants. We use fuzzy equivalence relations to model position specificity to develop a generalization of position-specific sequence kernels. This is not just a fascinating link, but it also leads to the creation of new kernels, such as the Elin kernel. We discovered that these kernels enable an explicit representation, allowing for improved computing efficiency and feature extraction.