The challenges of explainable AI in biomedical data science
Henry Han, Xiangrong Liu · BMC Bioinformatics · 2021
With the surge of biomedical data science, more and more AI techniques are employed to discover knowledge, unveil latent data behavior, generate new insight, and seek optimal strategies in decision making.Different AI methods have been proposed and developed in almost all different biomedical data science fields that range from drug discovery, electronic medical records (EMRs) data automation, single-cell RNA sequencing, early disease diagnosis, COVID research, and healthcare analytics.The AI methods and systems also generate a massive amount of data or big data that not only bring unpreceded progress in biomedical fields but also new challenges for AI.One of the key challenges should be the explainability of AI in biomedical data science problem-solving.It refers to that an AI method or system should not only bring good results but also have good interpretability, i.e., let users know why this way is the optimal one rather than the others.The existing AI methods employed in biomedical data science generally lack good explainability and may not create trustworthiness and transparency in usage well.For example, a deep learning model may bring good accuracy in disease diagnosis by analyzing corresponding bioimages, but it can be hard to explain well about the setting of thousands of parameters in the model.It can be possible that some small perturbations of the parameters may generate totally different learning results and challenge the robustness and stability of the deep learning model.Since AI models cannot explain themselves well, it is likely to encounter a high risk to make an incorrect decision making and decrease its trustworthiness and reliability, even if it has the advantage in accuracy, speed, or complicate data relationship revealing.On the other hand, the AI interpretation issue has been raised almost ten years ago in some subfield of biomedical data science such as bioinformatics.For instance, bioinformaticians found that the gene markers or network markers recommended from an AI disease diagnosis system may not explain themselves, i.e., the identified markers not only cannot apply themselves well in clinical practice, but also those markers that do well in the clinical practice may not be recommended from the AI system [1].