Accuracy Improvement by Occurrence Probability of Service Identification based on SNI
Ryo Asaoka, Akihiro Nakao, Masato Oguchi, Saneyasu Yamaguchi · 2022
Identifying the service from IP flows is beneficial for a variety of purposes, such as traffic priority control or achieving zero-rating services. Identification methods based on IP addresses and port numbers have limitations in accuracy, and methods based on inspection of packet payloads have been discussed. Since many recent packets are encrypted with TLS, it is important to study a method that can be applied to packets encrypted with TLS. As a method applicable for TLS encrypted data, a method based on the occurring SNI and Bayesian inference was proposed. This method searches the previously investigated occurrence SNI database for the SNI occurrence vector that perfectly matches the SNI occurrence vector of the IP flow to be identified. This method gives up identification if there is no occurrence SNI vector pair that perfectly matches. This giving up identification is considered to decrease its accuracy. In this paper, we propose a method for improving the accuracy by excluding SNIs that are not important for identification. This exclusion reduces giving up identification. We consider SNIs that have a low occurrence probability and SNIs that occur only in the traffic to be identified as non-important SNIs, and identify the service without these non-important SNIs. Our evaluation shows that the proposed method can achieve higher identification accuracy than existing methods.