Binary Classification is Enough: A Lightweight Strategy for Drug Screening with Small Datasets
Wu Liang, Xiaoguang Ma · 2023
Drug screening is an extremely costly and time-consuming process, wherein only small datasets are available in practice. We presented a particular method to estimate values of inhibition constant(Ki) or half-maximal inhibition concentration(IC50) of unknown compounds through a lightweight mutual information and logistic regression(MI-LR) united model that only needed to be trained on a small dataset. Biologists could then use this model to determine whether the compounds were initially eligible for screening, increasing efficiency of their work. A data augmentation strategy was used to sort independent samples of training datasets and solved the problem of sample shortage caused by the lightweight model, and transform a prediction task into a simpler binary classification task. In addition, we proposed an effective constraint mechanism to deal with the case when the classification results were contrary to the facts. By accurately predicting the interval of its inhibitory effect, we can improve the efficiency and accuracy of drug screening. Numerous evaluations on the Ki and IC50 dataset demonstrated high reliability of the MI-LR united approach to sort compounds according to a selected set of molecular descriptors.