Review of Large Language Models for Genomic Data and Medical Text
Devansh Sharma, Suraiya Jabin · International Journal of Bioinformatics and Intelligent Computing · 2025
With introduction of Transformer model in 2017 by a team of researchers at Google Brain, the field of Natural Language Processing was totally revolutionized. Google Translate started translating between two languages with more and more accuracy, as it was released from the clutches of legacy method of statistical machine translation and upgraded with Transformer model. Soon, these models were extended to other domains such as computer vision and time series data analysis. At the same time, the capabilities of these models were extended for a variety of genomic data for example whole genome sequences and protein sequences, the field of genomic data analysis was freed from Kmer count based hand-crafted features to sophisticated semantic capturing embeddings which were obtained with training of Transformer model using genomic data for certain biological tasks at hand for example enhancer prediction on epigenomics data or disease diagnosis using multi-omics data. This paper attempts to review and interpret the most recent large language models specially designed and trained for interpreting the semantics of whole genome sequence data and the medical text.