Crawling and Analysis of Dark Network Data
Ying Yang, Guichun Zhu, Lina Yang, Huanhuan Yu · 2020
Due to its anonymity and non-traceability, it is very difficult to research websites on the dark network. The research of the dark network is very important for our network security. Now there is very little data for studying the dark network, so we independently developed dark web crawler that runs automatically. This article will detail the implementation process of our dark web crawler and the data analysis process of crawled data. Currently, we can use crawled data to detect if multiple urls belong to the same site. We can use data to extract features of similar websites and we have generated an ever-increaing data set that can be used for simple website classification.We use the crawled data as a categorical dataset to categorize newly discovered urls.When we get the a certain number of new urls, we crawl again and the crawled data will be added to the previous data set. After multiple rounds of crawling, our data sets will be more and more abundant. through our approach, we can solve the problem that the dark network data is small, researchers can use our method to get enough data to study all aspects of the dark network.