SHIFA: SBERT-Based Healthcare Information Focused Arabic Question Answering
Rahaf Alruwaithi, Sarah Omar Alhumoud · IEEE Access · 2025
Question Answering (QA) systems have been developed as a promising solution and an efficient approach to retrieving significant information over the Internet. The answer selection is one of the key components of many types of QA applications, which requires addressing a semantic gap between question-answer pairs. The idea behind this research is to use the SBERT method for entities and context patterns to effectively capture the semantics betweenQApairs in two differentways: the feature extractor and the transfer learning model. On the one hand, we employ the Arabic BERT model with SBERT, fine-tuning its parameters on the AraMed datasets to transfer its knowledge for Arabic answer selection classification. On the other hand, we inquire about SBERT performance as a feature extractor model by combining it with LSTM, GRU and CNN classifiers. There are three models presented in SHIFA: the HT-SBERT-LSTM, HT-SBERT-GRU and HT-SBERT-CNN models deal with the problem of representing contextual data. The results show that the fine-tuned AraBERTv0.2 with the SBERT model accomplishes state-of-the-art performance results and attains up to 87% in terms of F1-score and accuracy compared to other Arabic pre-trained language models such as AraBERTv2, CAMeLBERT, mBERT, and LaBSE. Besides, the HT-SBERT-LSTM reaches an accuracy of 94.25%, while the HT-SBERT-GRU reaches an accuracy of 93.70%. The SHIFA models show that mixing sentence-level embeddings with sequence models achieves competitively contextual and semantic representations.