A Scalable Computer Vision Framework For Mobile Device Auto-Typing
Zongyi Joe Liu, Rohit Kumar Gupta, Kun Chen, Bruce Ferry, Simon Lacasse · 2020
In this paper, we present a computer vision framework that controls robots to auto type on a mobile device such as an Android phone or an iPad. The framework consists of three parts: (i) an image undistortion and segmentation algorithm that supports images captured by a top mounted camera or a side mounted camera, (ii) a deep neural network (DNN) algorithm that detects the keyboard region and recognizes isolated characters from an input image, and (iii) a grid click algorithm to construct data points to compute the image to robot coordinate translation matrix that is scalable to any touch device of different sizes. We have demonstrated in this paper the accuracy and scalability of the proposed system.