Using Convolutional Neural Networks to Identify COVID-19 Infection from Lymphocyte Morphology in Digital Optical Micrographs
Daniel Walker, Richard Ling, Robert Banthorpe, Mahesh Prahladan, Nicholas H. M. Caldwell · Microscopy and Microanalysis · 2025
Lymphopenia (absolute lymphocyte count of <1x109/L) with lymphoplasmacytic morphology, is one of the clinical manifestations of SARS-CoV-2 aka COVID-19 e.g. [1]. During the early pandemic, clinician colleagues noted that the lymphocytes from COVID-19 patients showed several striking changes compared to the normal lymphocytes and lymphocytes observed in other viral infections. Other anecdotal reports also noted similar abnormalities such as lymphoplasmacytoid lymphocytes with an eccentric nucleus, deeply basophilic cytoplasm and a prominent paranuclear hof e.g. [2]. Morphological features of interest, including cell nuclear-cytoplasmic (NC) ratio, shape of the nucleus, chromatin pattern (fine or coarse, speckled, open, condensed), presence and numbers of nucleolus, presence of cytoplasmic vesicles/granules cytoplasmic pseudopodia and presence of red cell scalloping, were all considered as the focus of investigation to be promising candidates as diagnostic markers, either individually or collectively, of COVID-19 infection. In parallel with additional clinical observations, it was hypothesised that a deep learning model could be developed that would be able to distinguish between the abnormal lymphocytes associated with a COVID-19 infected patient and normal lymphocytes from a healthy patient. The dataset for the study were high quality images, processed by digital scanning of blood smears, acquired using a CellaVision ® DM96 Sysmex single, standalone haematology analyser. In more detail, the lymphocytes were obtained from a venous sample collected in EDTA type anticoagulant and stained on an automated tracker using SP-100 Sysmex analyser. The staining programme used conventional staining of haematoxylin and eosin (H&E) staining with the reagents as follows: May Grunwald (MG) and Geimsa (Biolyon, Frnace) (MG pure time: 2.5min, MG dilute time:3min, Giemsa time: 7min, rinse 0 min and drying time: 5min). The objective magnification was x100 using oil extracted on a camera with automated 200 cell differential made using the CellaVision blood differential software. For initial experimentation and to prove the viability of deep learning models, the team used a high-end computing workstation with 7.3 TB storage, 32 GB RAM, Intel 8600k 6-core 6-thread CPU, and a GPU Nvidia Pascal 1080 card with 8 GB VRAM. For rapid exploration of a wider configuration space of the deep learning models, the AI Compute Server at the University of Suffolk was used. The Compute Server has 159 TB storage, 512 GB RAM, an AMD Epyc 7542 32-core 64-thread CPU, and 4 Nvidia Ampere 100 GPUs with 40 GB RAM. Python and TensorFlow.Keras were used for development. The VGG-16 and AlexNet architectures have both been used in Deep Learning research to classify leukocytes with very promising results, suggesting that these model architectures may be suitable for use in lymphocyte research [3, 4]. As the image data at 363x363x3 was larger than the image data sizes expected by these pre-existing models, the model design drew inspiration from their arrangements of layers and other architectural choices rather than attempting to leverage them as pre-trained models. To increase the quantity of images available for training and testing the models, a variety of data augmentation techniques were employed to create simulated images by morphing real images in controlled ways, e.g. shifting the image in either axis, flipping the image horizontally or vertically, zooming, shearing, rotating and brightness changes. This can reduce the risk of overfitting and improve accuracy [5, 6]. The first phase of the study sought to classify lymphocytes into two categories – “activated” where the lymphocytes exhibited the abnormal features associated with COVID-19 and “normal” where the lymphocytes had normal characteristics (and came from healthy patients). A series of architectures were designed and trialled drawing upon the VGG-16 architecture as a baseline where the model was arranged as a set of blocks of convolutional layers (applying an operation akin to kernels in conventional image processing) and maxpooling layers (which downsamples the output of the preceding group of convolutional layers to select a maximum value within a local region). To reduce the computational requirements, each block consisted of two convolution layers and one maxpooling layer, with each block working on a progressively smaller image. The final block consisted of a fully connected layer (using 512 neurons rather than 4096 as per VGG-16), a dropout layer (which ignores some randomly chosen neurons on each training cycle to prevent overfitting) and a final output layer. The dataset had 392 training images (200 normal, 192 activated), 243 validation images (121 normal, 123 activated), and 123 images used for testing (61 normal, 62 activated). Multiple variants of this model (with additional dropout layers in blocks before or after the maxpooling layer), with different learning rates were tested to narrow in on the best configuration. The optimal configuration achieved a balanced accuracy of 98.37% (misidentifying one normal and one activated image). The second phase of the study sought to classify lymphocytes into three categories – “COVID-positive”, “COVID-negative” (healthy patients) and infected with EBV (Epstein-Barr-Virus). Epstein-Barr-Virus was chosen due to the virus again leading to abnormal lymphocytes. A much larger dataset was available of 7416 training images (2384 EBV, 2491 negative, 2541 positive), 2099 validation (699 EBV, 700 negative, 700 positive) and a set of 1050 images (350 each category) used only for testing. The existing optimal model configuration was adjusted to permit three outcomes, but it overfitted constantly and was abandoned. The technique of batch normalisation (where inputs to a layer are normalised) was used to stabilise the training [7]. A new variant model (of up to 29 layers) with batch normalisation and dropout layers in each block was created. 512 configurations were tested to identify viable learning rate, followed by another 480 configurations. Despite this, the leading candidate configurations were only able to reach a balanced accuracy in the range of 80%-87%, which compares unfavourably with the 98.6% accuracy achieved elsewhere, albeit using a much larger 71-layer Xception architecture [8]. In conclusion, although deep learning approaches can achieve high accuracy, success on one problem is not certain to be transferrable and the potential solution space in terms of number of layers, types and sizes of layers, and many other parameters makes finding effective and computationally efficient (so arguably more sustainable) model configurations a difficult task [9].