Depth-induced prompt learning for laparoscopic liver landmark detection

Ruize Cui, Weixin Si, Zhixi Li, Kai Wang, Jialun Pei, Pheng‐Ann Heng, Jing Qin · Medical Image Analysis · 2026

• A new liver landmark detection dataset, L3D-2K, comprising 2,000 keyframes sourced from surgical videos with professional annotations. • A novel deep learning framework D2GPLand+ that utilizes RGB-D information for laparoscopic liver landmark detections. • Proposing the DPE module, which incorporates learnable prompts with contrastive learning to discriminate the geometric features of different landmark categories from depth clues. • Introducing the CUMamba block that concurrently conducts cross-modal interactions on spatial dimension and feature reparameterization on channel dimension for effective RGB-D fusion. • Introducing the AFA scheme to highlight anatomical structures by implicit and explicit edge emphasis and controlling detail levels. Laparoscopic liver surgery presents a highly intricate intraoperative environment with significant liver deformation, posing challenges for surgeons in locating critical liver structures. Anatomical liver landmarks can greatly assist surgeons in spatial perception in laparoscopic scenarios and facilitate preoperative-to-intraoperative registration. To advance research in liver landmark detection, we develop a new dataset called L3D-2K , comprising 2,000 keyframes with expert landmark annotations from surgical videos of 47 patients. Accordingly, we propose a baseline, D 2 GPLand+, which effectively leverages depth modality to boost landmark detection performance. Concretely, we introduce a Depth-aware Prompt Embedding (DPE) scheme, which dynamically extracts class-related global geometric cues with the guidance of self-supervised prompts from the SAM encoder. Further, a Cross-dimension Unified Mamba (CUMamba) block is designed to comprehensively incorporate RGB and depth features with the concurrent spatial and channel scanning mechanism. Besides, we bring out an Anatomical Feature Augmentation (AFA) module that captures anatomical cues and emphasizes key structures by optimizing feature granularity. For benchmarking purposes, we evaluate our method and 17 mainstream detection models on L3D, L3D-2K, and P2ILF datasets. Experimental results demonstrate that D 2 GPLand+ obtains superior performance on all three datasets. Our approach provides surgeons with guiding clues that facilitate surgical operations and decision-making in complex laparoscopic surgery. Our code and dataset are available at https://github.com/cuiruize/D2GPLand-Plus .

Read the paper · More papers on PaperTik