SFNet: Synergistic Fusion Network for Text Localization in Natural Scene Images
Suman Suman, H N Champa · 2024
Prior methodologies for the detection of scene text have demonstrated promising outcomes across many evaluation metrics. Nevertheless, despite the utilization of deep neural network models, their effectiveness is sometimes limited in complicated situations due to the intricate interplay of several stages and components within the pipelines, which ultimately governs overall performance. This study presents a straightforward yet robust methodology for identifying textual content in practical environments. Phrases or text lines may be predicted in images of varying orientations and quadrilateral forms using the suggested pipeline’s single neural network. This eliminates the need for superfluous intermediary steps such as candidate aggregation and word division. Due to the streamlined nature of our SFNet(Synergistic Fusion Network), our primary emphasis is directed at the development of loss functions and neural network design. The suggested approach has been evaluated using a well-established dataset including ICDAR 2015. The results of these experiments demonstrate that the new technique outperforms existing techniques in terms of both accuracy and efficiency. The proposed methodology demonstrates a performance metric, namely an F-score of 0.89 while operating at a frame rate of 13.2 frames per second and a resolution of 720p using the ICDAR 2015 dataset.