MyoLA: Classifying Myocardial Infarction Based on Transcription Texts Leveraging AI and Linguistic Algorithms
Ankna Litoriya, Abhishek Shrivastava, Santosh Kumar · 2024
Myocardial infarction (MI), also known as heart attack, occurs when blood flow to the heart is obstructed, causing damage to the heart muscle. Diagnosing MI remains challenging due to variable symptoms, inconsistent data, incomplete patient records, and analysis of clinical information like test symptoms, reports, and analysis is challenging for medical professionals. To address these challenges, we introduce MyoLA, a novel framework that combines Linguistic Algorithms (LA) with AI techniques for MI classification from medical transcriptions. MyoLA integrates natural language processing (NLP), principal component analysis (PCA), logistic regression (LR), and random forest models to enhance real-time MI detection and diagnostic accuracy. Utilized a diverse dataset from mtsamples.com, consisting of transcriptions from various medical specialties. The preprocessing pipeline involved text cleaning, lemmatization, and TF-IDF for feature extraction, while class imbalance was handled using the SMOTETomek technique. Dimensionality reduction was achieved through t-distributed stochastic neighbor embedding (t-SNE). Multiple models, including transformer-based models like BERT, were evaluated. Logistic Regression with SMOTETomek yielded an F1-score of 0.94, while BERT achieved superior performance with an F1-score of 0.91 and an accuracy of 98.2%, albeit with higher computational costs. Visualization through confusion matrices and t-SNE plots demonstrated significant improvements in classification accuracy. These findings confirm the potential of integrating advanced NLP techniques and AI models to improve the accuracy and efficiency of MI diagnosis from medical transcriptions.