Symbol Recognition Method for Railway Catenary Layout Drawings Based on Deep Learning
Qi Sun, Mengxin Zhu, Minzhi Li, Gaoju Li, Weizhi Deng · Symmetry · 2025
Railway catenary layout drawings (RCLDs) have the characteristics of upper and lower symmetry, a large drawing size, a small size, high similarity among target symbols, and an uneven distribution of symbol categories. These factors make the symbol detection task more complex and challenging. To address the aforementioned challenges, this paper proposes three enhancements to YOLOv8n to improve symbol detection performance and integrates an improved denoising diffusion probabilistic model (IDDPM) to mitigate the imbalance in symbol category distribution. First, the multi-scale dilated attention (MSDA) is introduced in the Neck part to enhance the model’s perception of the global context in complex RCLD scenes, so that it can more effectively capture the symbol information distributed in different scales and backgrounds. Secondly, the receptive field attention convolution (RFAConv) is used in the detection head to replace the standard convolution, to improve the ability to focus on the target symbols in RCLDs and effectively alleviate the occlusion interference between symbols. Finally, the dynamic upsampler (DySample) is used to enhance the clarity and positioning accuracy of the edge area of small target symbols in RCLDs and enhance the detection of small targets. The above design made targeted optimizations to resolve the problems of symbol and background interference, character overlap, and symbol category imbalances in complex scenes in RCLDs, effectively improving the overall detection performance of the model. Compared with the baseline YOLOv8n model, the improved YOLOv8n achieves increases of 2.9% in F1, 1.9% in [email protected], and 1.7% in [email protected]:0.95. With the introduction of synthetic data, the recognition of minority-class symbols is further enhanced, leading to additional gains of 4%, 3.8%, and 14% in F1, [email protected], and [email protected]:0.95, respectively. These results demonstrate the effectiveness and superiority of the proposed method in improving detection performance.