Performance Analysis Using Categorical Boosting Classifier (CatBoost), Ensemble Hard Voting Classifier (EHVC) & Convolution Neural Network (CNN) for DNA Sequence Classification
Falguni Adhikary, Swarup Sarkar, Himanshu Pal · 2025
Recently, Machine Learning (ML) has become an essential methodology for medical science. Deoxyribonucleic Acid (DNA) carries all the genetic information of every individual. The extraction of features from DNA genetic variations & identifying species are quite difficult without using ML and deep learning as it can learn the features automatically. It cannot process the raw data as DNA sequences are represented as the combination of adenine (A), thymine (T), guanine (G) & cytosine (C). A conversion of string data to a tokenized vector is needed as it mapped with nucleotides to a numerical token. A k-mer substring of length k hexamer words of size 6 is used here as its coming from a longer DNA sequence with a combination of ATGC. In this paper, the Categorical Boosting classifier (CatBoost), an ensemble algorithm with hard voting, is proposed with the combination of K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Random Forest (RF) model together to make better accuracy and robustness of the DNA classification task after vectorization of k hexamer of 6 word sizes. These enable the correctness of medicine, early disease diagnosis, and genomic observation of infectious diseases. The proposed model also integrates the deep learning model, which consists of convolutional neural networks (CNNs), to improve the accuracy of DNA sequence classification over the CatBoost and the EHVC. Our model accuracy of the ensemble method using hard voting achieved 93% and CNN 95%.