Neural networks may outperform classical regressions, but only when non-linear relationships are considered

Carlos Hernandez-Vaquero, Daniel Hernández‐Vaquero · European Journal of Cardio-Thoracic Surgery · 2021

We have read with great interest and admiration the article by Benedetto et al. [1] about the comparison of modern machine learning with traditional regression models for the prediction of mortality after cardiac surgery. This issue is extremely interesting and novel due to the intensity in which machine learning has emerged in medicine and many other disciplines [2]. After analyzing >28 000 patients, they concluded that neural networks do not provide an advantage over the usual logistic regression. The limited number of events, the inability to perform ML model hyper-tuning and the absence of continuous variables were identified as possible reasons for this lack of improvement [1]. The limited number of events may be a limitation of modern machine learning techniques, but this is also the case for logistic regression models. Something similar occurs with the absence of continuous variables. This may be a disadvantage for neural networks, but the loss of information that occurs when a variable is categorized has also a negative effect on any regression and even on any statistical method. It is perfectly possible that a logistic regression provides very good results and neural networks do not improve them. This usually happens when the input variables investigated to predict an outcome are related linearly, which is an assumption that needs to be met when using logistic regressions [3]. Doctors use those variables that proved to be predictors in logistic regression models, such as EuroScore or STS risk score. This is, variables that showed to have an understandable, straightforward and therefore linear relationship with the outcome. The benefits of using neural networks with respect to classical predictors emerge when non-linear relationships are present in the input and output variables [4]. These relationships are hidden to the human eye and therefore not generally selected for developing classical predictors. The focus of future work should involve finding the appropriate input variables to profit from using neural networks, rather than limiting the study to using the same features, which are known to work with linear models. In addition, from our point of view, randomness, the surgical team or unpredictable variables can play a key role on the risk of mortality after cardiac surgery. Machine learning will not overcome the possibility of randomness or an extraordinary surgical team, but this technique could overcome previous regressions when adequate variables with non-linear relationships are used.

Read the paper · More papers on PaperTik