ArthroNet: a monocular depth estimation technique with 3D segmented maps for knee arthroscopy
Shahnewaz Ali, Ajay Kumar Pandey · Intelligent Medicine · 2022
: Lack of depth perception from medical imaging systems is one of the long-standing technological limitations of minimally invasive surgeries. The ability to visualize anatomical structures in 3D can improve conventional arthroscopic surgeries, as a full 3D semantic representation of the surgical site can directly improve surgeons’ ability. It also brings the possibility of intraoperative image registration with pre-operative clinical records for the development of semi-autonomous, and fully autonomous platforms. : Depth estimation and segmentation processes of feature- and texture-less tissue structures is an extremely challenging task. The lack of accurate ground-truth depth data, non-even lighting conditions with the occurrences of intra-frame over- and under-exposed regions, and occlusions limit the application of stereo vision and monocular supervised learning techniques inside the joint space. On the other hand, unsupervised/self-supervised monocular depth estimation techniques suffer from depth discontinuity, poor depth estimation in texture-less regions, no association with depth scale, and poor depth gradients are common. The provision of fully segmented 3D maps solves the grand visualization challenge of knee arthroscopy, and our method is widely applicable to other forms of minimally invasive surgeries and 3D reconstruction of medical images in general. We apply a novel technique that provides the ability to combine both supervised and self-supervised loss terms, and in doing so eliminates the drawback of each technique. It enables the estimation of edge-preserving depth maps from a single untextured arthroscopic frame. The proposed image acquisition technique projects artificial textures on the surface to improve the quality of disparity maps from stereo images. Moreover, integration of attention-ware multi-scale feature extraction technique along with scene global contextual constraints and multiscale depth fusion, the model able to predict reliable and accurate tissue depth of the surgical sites that complies with scene geometry. : A total of 4128 stereo frames from a knee phantom were used to train a network, and during the pre-trained stage, the network learns disparity maps from the stereo images. The fine-tuned training phase uses 12,695 knee arthroscopic stereo frames from cadaver experiments along with their corresponding coarse disparity maps obtained from the stereo matching technique. In a supervised fashion, the network learns the left image to the disparity map transformation process, whereas the self-supervised loss term refines the coarse depth map by minimizing reprojection, gradients, and structural dissimilarity loss. Together, our method produces high-quality 3D maps with minimum re-projection loss that are 0.0004132 (structural similarity index), 0.00036120156 (L1 error distance) and 6.591908e-05 (L1 gradient error distance). : Machine learning techniques for monocular depth prediction is studied to infer accurate depth maps from a single-color arthroscopic video frame. Moreover, the study integrates segmentation model hence, 3D segmented maps were inferred that provides extended perception ability and tissue awareness.