Implementation of Synthetic Minority Over-Sampling Technique in the Anaemia Classification Using the LSTM and Bi-LSTM Algorithms
Yuri Pamungkas, Ratri Dwi Indriani, Zain Budi Syulthoni · 2024
Anaemiais a health disorder characterized by a lack of red blood cells in a person's body. Anaemia sufferers will tire more quickly, and their faces look paler than normal people's. Anaemia can occur over a short or long period (depending on the severity). If anaemia continues to be ignored, complications can worsen a person's health condition. Considering the impact caused by anaemia, a solution is needed to detect anaemia precisely and accurately. One way can be taken is by utilizing an artificial intelligence-based system to detect or classify anaemia based on its symptoms. Therefore, we tried to analyze factors related to anaemia in this study and carry out anaemia classification based on the LSTM and Bi-LSTM algorithms. The dataset used in this study came from the Kaggle repository, which contains medical record information for 1421 patients (620 patients diagnosed with anaemia and 801 patients with non-anaemia). The patient's medical record information includes gender, blood haemoglobin level, MCH, MCHC, MCV, and diagnosis results. In the research dataset, we also applied SMOTE to balance the data classes for anaemia and non-anaemia sufferers and compare the classification results' performance. Based on the research results, the haemoglobin level factor has the highest correlation value of 0.8 compared to other factors such as gender (0.25), MCHC (0.05), MCH (0.03), and MCV (0.02). Meanwhile, the classification results show that the use of SMOTE can increase the specificity (100%), precision (100%), F1-score (98.93%), and accuracy (98.86%) of the LSTM algorithm during the classification process.