Exploring the effect of observer task in evaluating performance of FFDM with in silico modeling
Dan Li, Andrey V. Makeev, Stephen J. Glick · 2024
Virtual clinical trials (VCTs)and physical phantom studies can be helpful in the evaluation of new breast imaging technology.1, 2 These methods are usually designed to provide an objective assessment of image quality by evaluating task-based performance from an observer (e.g., human or model observer). Previous VCTs and phantom studies typically use a detection task (often a signal-known exactly task) to assess performance. However, real clinical studies are usually more complicated than the simple task of lesion detection. Most breast imaging clinical studies reported in the literature involve the observer both detecting whether a suspicious lesion is present, as well as rating how likely it is that the lesion is malignant. In fact, the decision variable commonly used in clinical breast imaging studies to form a receiver operating characteristic (ROC) curve is the probability of malignancy (POM). Designing VCTs and physical phantom studies that accurately model clinical studies is challenging because it is difficult to model malignant and benign lesions. One lesion feature that that is commonly used to discriminate malignant and benign lesions is the border of the lesion, with malignant lesions typically portraying spiculations. To more accurately model clinical breast imaging studies, we explore the use of a classification task instead of a detection task, that evaluates the reader’s performance in discriminating between spiculated and lobular (non-spiculated) masses. To assess this classification task, we use a convolutional neural network (CNN) model trained to differentiate spiculated and non-spiculated masses using Monte Carlo simulated images. The basis for using CNN observer models is that a previous study showed that the CNN based framework can be trained to approximate the ideal observer.3 In addition to the mass classification task, we also examine this CNN model observer in a more traditional detection task. This study evaluates the effect of dose and resolution and shows that different conclusions might result depending on whether the VCT model evaluates a mass detection task or a mass classification task.