NEEDLE: Nurse Education Enhanced by Vision-based Deep Learning Evaluation

Matthias Tschöpe, Stefan Gerd Fritsch, Vítor Fortes Rey, Niranjan Narendra Nandurkar, Sarah Trevenna, Eloise Monger, Paul Lukowicz · 2025

Training nurses in procedures such as venipuncture and cannulation is time-consuming and requires a teacher to supervise and provide verbal feedback. Automating this process could allow students to practice independently, reducing the need for constant supervision. Recent advances in vision-based deep learning models offer the ability to classify and evaluate students’ performance in video recordings, while a Large Language Model can provide feedback. This work lays the foundations for such a system by comparing the performance of six state-of-the-art video classification models to classify key activities of venipuncture and cannulation sessions recorded in a teaching hospital. We also evaluate the zero-shot feasibility of the vision-language model Qwen2-VL (2B, 7B, 72B parameters). The performance is evaluated based on the macro $F_{1}$-Score, VRAM utilization, and energy consumption. For cannulation, the Swin3D base model ($\mathbf{8 8 M}$ parameters) achieves a macro $F_{1}$-Score of $\mathbf{5 6. 7 9 \%}$, while the largest Qwen2VL model achieves only $27.67 \%$. A similar trend is observed for venipuncture ($\mathbf{4 3. 7 1 \%}$ vs. $\mathbf{2 3. 2 3 \%}$). The Swin3D base model is also more energy-efficient, consuming 15 to 35 times less energy than Qwen2-VL.

Read the paper · More papers on PaperTik