Evaluating Model Compression Techniques for On-Device Browser Sign Language Recognition Inference
Mauricio Salazar, Danushka Bandara, Gissella Bejarano · 2024
Due to the improvement of sign language (SL) recognition models, more work has started to build upon them to create end-user platforms such as video-based dictionaries. Although most of them base their design on cloud services or APIs, this approach can become expensive, especially when working with such substantial dimensional data as videos. Other realities, such as regions with low internet connection and few economic resources, can benefit from the offline alternative for deploying these systems. However, models deployed in websites' front-end or edge devices have some constraints on their size. Therefore, our work evaluates two model compression techniques to enhance the performance of on-device inference for a SL recognition model. We test knowledge quantization and distillation on three datasets varied in size, ultimately deploying and evaluating them on a locally hosted web page. To the best of our knowledge, this is the first evaluation of a quantized and distilled transformer-based SL recognition model for on-device browser inference. Furthermore, the combination of both compression techniques achieves a size reduction of up to 17.83 times and a speedup of 5.9 times the baseline.