ViTDFNN: A Vision Transformer Enabled Deep Fuzzy Neural Network for Detecting Sleep Apnea-Hypopnea Syndrome in the Internet of Medical Things

Na Ying, Hongyu Li, Zhi Zhang, Yong Zhou, Huahua Chen, Meng Lin Yang · IEEE Transactions on Fuzzy Systems · 2024

Sleep apnea-hypopnea syndrome (SAHS) is a disease that seriously affects human sleep. Due to its strong concealment and significant harm, it has gradually attracted people's attention. Therefore, it is necessary to use the SAHS detection model to screen the disease as early as possible. However, most of the current SAHS detection models use convolutional neural network (CNN) architectures such as Resnet, which are ineffective in capturing images' details. Moreover, the Vision Transformer (ViT) model relies heavily on the number of SaO2 signal image data during training. Otherwise, it isn't easy to achieve better performance. In addition, due to the existence of fuzzy and uncertain information in the SaO2 signal image, the detection accuracy of previous models is low. Therefore, in this work, we propose a vision transformer enabled deep fuzzy neural network (ViTDFNN) to detect SAHS in an Internet of Medical Things platform. The ViTDFNN model can analyze the patient's vital signs and sleep status on time and identify potential SAHS conditions. The ViTDFNN model first uses ViT to learn the features of the SaO2 signal image so that global information and long-range dependencies can be learned. Subsequently, the ViTDFNN model uses VGG to extract the features of the SaO2 signal image to learn the local features in the image, which is convenient for distinguishing the texture, edge, and other details. Finally, the ViTDFNN model linearly fuses the two image features and inputs them into the deep fuzzy neural network and fully connected network to detect the patient's symptoms. The deep fuzzy neural network can convert each deterministic pixel into an uncertain pixel, thereby solving the fuzzy and uncertain information in the SaO2 signal image. In addition, the ViTDFNN model also uses the masked strategy to form a task of image reconstruction, which is used to reduce the dependence of the ViT model on the labeled SaO2 signal image. We verified the superiority of ViTDFNN by conducting extensive experiments on a real sleep dataset. The scores of the ViTDFNN model under the F1-Score, Recall and Accuracy metrics are 0.994, 0.992 and 0.994, respectively.

Read the paper · More papers on PaperTik