Rapid Homoglyph Prediction and Detection
Avi Ginsberg, Yu Juan Cui · 2018
As technology permeates the globe, it is becoming necessary to support an increasing number of languages, and hence character sets. Many character sets contain characters that look similar or identical to characters from other character sets (known as homoglyph characters). This poses a number of unique problems, such as the homoglyph URL attacks frequently used in phishing scams. The previously proposed solutions to these problems have mostly focused on restricting the characters available to users. This paper proposes a new approach which can predict and detect homoglyphs with a high level of accuracy. This methodology simulates the human behavior of "visual scanning" and allows large amounts of visual data to be represented in a format that is easily and quickly searchable. Additionally, it employs a pre-processing time-memory trade-off to enable the detection and prediction of homoglyphs in fractions of a second.