AN ALGORITHM FOR GENERATING WORD LISTS WITH A SPECIFIED BIGRAM DISTRIBUTION FOR KEYSTROKE DYNAMICS BIOMETRIC TEMPLATE REGISTRATION
Evgeniy V. Shklyar · Bezopasnost informacionnyh tehnology · 2025
The paper proposes a solution to the problem of ensuring the possibility of obtaining identical sets of bigrams contained in different word sets. The approach is based on a text data generation algorithm designed to provide stable characteristics of biometric parameters when entering free text. The proposed solution is implemented as a three-stage process: selection of word forms from the modern Russian language, construction of bigram distributions, and formation of vocabulary lists. The effectiveness of the proposed method is experimentally confirmed. Optimal parameters for the registration of keystroke dynamics samples have been determined and empirically validated: the number of words — 32, the length of each word — from 5 to 7 characters, and the total length of the sequence — 204 characters. These parameters ensure a registration text input time of less than one minute at an average typing speed of 220 characters per minute. The reliability of the results is supported by a high degree of similarity between the bigram distributions of the test and reference word sets — up to 94.18%. This provides a false acceptance rate (FAR) and a false rejection rate (FRR) that meet regulatory requirements. The results confirm the applicability of the proposed algorithm in biometric authentication and user identification systems, including conditions of free-text input.