A Machine Learning algorithm for Predicting Translation Initiation Start Positions
Stefanos Digenis, Dimitris Grigoriadis, Marios Miliotis, Artemis G. Hatzigeorgiou · 2024
Translation Initiation Start (TIS) sites are specific regions in genes, where the protein synthesis starts. In this study, the specific problem of RNA sequence classification is addressed through the deployment of convolutional neural networks (CNNs) employing features extracted from genomic sequences. Translation initiation is a crucial step in protein synthesis, indicating the sites where the translation mechanism takes place and such, accurate prediction of its start sites is a significant step towards deeper understanding and harnessing of its complex characteristics. To this aim, a state-of-the-art deep learning model has been developed which predicts the presence/absence of translation initiation start sites within RNA sequences. The proposed deep learning model leverages the power of CNNs to effectively capture relevant features in RNA sequences. Trained on a comprehensive dataset of manually annotated translation initiation sites, the model learns to discern key patterns and motifs associated with initiation starts. Through rigorous evaluation, a model comprising sequence, conservation and hexamer features exhibits a remarkably high accuracy rate of 93.75%. This research holds promise for various biological applications, including gene annotation, drug target identification, and understanding the mechanisms of protein synthesis. The ability to reliably predict translation initiation start sites using deep learning provides a valuable tool for computational biologists and experimental researchers alike.