WatchLogger: Keystroke Detection and Recognition of Typed Words Using Smartwatch
Gangkai Li, Yugo Nakamura, Hyuckjin Choi, Shogo Fukushima, Yutaka Arakawa · Sensors and Materials · 2024
Nowadays, more and more people are wearing smartwatches in their daily lives.The various sensors embedded in smartwatches bring the ability to evaluate users' status as well as the risk of privacy issues.For example, if users are typing on keyboards while wearing smartwatches, the attacker can know the typed contents from the sensor data collected by the malicious applications that are installed on the targets' smartwatches.In this paper, we propose WatchLogger, a framework using audio and accelerometer signals to recognize the English words being typed, to demonstrate how to implement the smartwatch-based side-channel attack.In contrast with previous studies that focused on the recognition of each key or pair of keys being pressed, WatchLogger aims to perform recognition on the scale of words.To achieve this goal, WatchLogger exploits the audio signals for segmentation and the accelerometer signals for classification.In addition, we propose an ensemble classification model to deal with the problem caused by too many words.Finally, we build the WTW-100 dataset (Wearable Typed Words dataset with 100 classes of words) using data from four participants and conduct experiments on the basis of this dataset.The experimental results show accuracies of 98.31 and 99.62% and F1 scores of 0.9745 and 0.9855 for keystroke detection and classification, respectively, and an accuracy of 79.76% for word classification, indicating a considerable performance of WatchLogger.Sensors and Materials, Vol.36, No. 10 (2024) Assuming that a user is typing using a keyboard while wearing a smartwatch, the attacker can clearly infer the typed contents from the sensor data of smart devices.(1)(2)(3) Intuitively, the typing event consists of three types of hand activity, namely, moving toward the key, pressing, and releasing.When typing different words, the trajectories of hand movements could be different, as well as the fingers used to press keys.These subtle differences are reflected in the unique patterns of wrist motions that can be measured by a wrist-worn accelerometer.Therefore, the typed contents can be inferred from accelerometer signals.In addition, the sounds of keystrokes provide extra information, such as whether someone is typing or not, and can be obtained from the microphone embedded in the smartwatch.Most previous works focus on the recognition of each key or pair of keys being pressed.Harrison et al. (4) proposed a framework that classifies 36 keystrokes on a MacBook keyboard on the basis of their sounds.Maiti et al. (1) examined the hand movement directions while typing and inferred the typed words from the series of directions.The advantage of these methods is that the number of target classes is small and fixed.However, the keystroke samples are hardly distinguishable owing to the short period of pressing one key (e.g., less than 1 s) and are vulnerable to external interference, such as environmental noise.Another problem is that using accelerometer signals to recognize each keystroke is nearly impossible because the trajectory of the hand when pressing one key depends not only on the current key but also on the previous key (e.g., the movements of pressing "e" in the words "are" and "be" are different), making it difficult to find a solution on the scale of a single key.In conclusion, the deficiency of key-by-key-based methods is that the information contained in pressing a single key is limited, making the system less robust.In this paper, instead of recognizing each key, we propose WatchLogger, a method that recognizes the scale of words using both audio and accelerometer signals.To enable word recognition, we propose three assumptions.First, only English words will be recognized.Second, all words in a sentence are segmented by the space key.Third, we only focus on the keys of lower-case a-z and the space key.Under these assumptions, WatchLogger is able to segment the signals into frames that represent each single word by finding all space keystrokes and can train a word classifier using the labeled frames.One problem is how to find space keystrokes.We note that the sound of pressing a space key is distinct from that of other keys on a keyboard, so we can recognize them by audio signals.In this paper, we propose a novel and efficient method to find all keystrokes in the audio signals and distinguish the space keystrokes, which solves several problems in traditional ways.In addition, the number of words is large in a real situation, making it difficult to classify all of the words.Actually, the attacker does not need to know every word being typed by the target user, that is, the keywords in the dictionary are enough.Considering that the number of keywords could still be large, we design an ensemble model that divides the word set into subsets to decrease the number of classes for each submodel.This work is an extended version of our previous research (5) and mainly has the following contributions: 1) We proposed a method to recognize typed words on the basis of audio and accelerometer signals.2) We proposed a model to recognize words on the word scale, different from the previously used scale of a single key or pair of keys.