Capture More Structured Context by Vision Transformers for Free-Hand Sketch Recognition
Weirong Guo, Sirui Liu, Yaoxiang Yu, Bo Cai · 2023
Free-hand sketch recognition presents unique challenges that require careful consideration of spatial and temporal properties. Convolutional neural networks, which are widely used for sketch recognition, have limited ability to extract structured context, leading to insufficient learning of spatial properties. In this study, we introduce Vision Transformers to the sketch recognition domain to address this issue. To the best of our knowledge, this is the first attempt to evaluate the applicability and effectiveness of Vision Transformers for sketch recognition. In addition, we also address the problem of semantically different categories appearing similar in sketches. To overcome this challenge, we propose a novel loss function for sketch recognition, termed Sketch Online Label Smoothing (SketchOLS). Our experiments on QuickDraw dataset demonstrate that the proposed approach is a robust and effective solution for sketch recognition. Remarkably, our method outperforms existing models even without the injection of additional temporal information. These results highlight the potential of our approach to enhance free-hand sketch recognition tasks.