A mathematical expression model of network asset data
Baolin He, Tao Yang, Xiaoyu Ma, Qingqing Chen · 2021
When the data sources of network assets collected by network detection technology need to use the traditional machine learning technology for research and analysis, it requires to carry out a comprehensive numerical conversion. However, this kind of data presents as a collection of various heterogeneous features and complex textual data. When numerical conversion is performed in order to be trained by machine learning algorithms, there is a lack of unified and comprehensive rules for the conversion of various heterogeneous features. At the same time, the numerical transformation of text features lacks rationality and validity. So the application of machine learning to network asset data is greatly restricted. In response to the above problems, this paper proposes a mathematical expression model for network asset data. Firstly, it defines the rules of timestamp, segment transfer method etc to convert various heterogeneous features, and then constructs the AWK (Adaptive-Word2vec-Kmeans) algorithm to convert the complex text data into numerical values. Based on the expression model, the data source obtained by Nmap is processed to a trainable data set. The full feasibility and effectiveness of the model are proved through experiments, and the converted data set is available and is more challenging.