Improving Tor Hidden Service Crawler Performance
Jong-Hyeon Park, Hyunsu Mun, Youngseok Lee · 2018
Tor is a popular anonymous network that provides the dark web called the hidden service not accessible with traditional web browsers. The number of hidden services using Tor's anonymous network has been steadily increasing every year with the demand of user privacy concerns as well as illegal services that abuse Tor's anonymity. As a result, how to collect and analyze the contents of Tor hidden service is not yet in full swing. In addition, due to the slow access speed of the Tor browser, it is difficult to observe the dynamics of the hidden services. In this work, we present an efficient Tor hidden service crawling method that can monitor the status of hidden services for the analysis of authentic hidden service contents. We have designed our Tor crawler with the virtualization software, Docker, and parallel Tor browsers to improve the crawling performance. In order to maximize the utilization, we analyzed the collected Tor hidden services by a clustering technique and reduced the crawling targets. Crawling time was 76% improved with 5 Docker containers with 5 browsers. From clustering same hidden services, we can also reduce 39% of crawling the Tor hidden services. From January 1, 2018 to May 31, 2018, we collected 25,261 onion addresses and found 2,527 unique hidden services. We crawled this hidden services twice a day, collect 456,739 HTML pages. We classified 722 hidden services contents and analyzed the Tor hidden service dynamics through collected HTML pages and crawler's log data. As a result, we found that the percentage of illegal services is higher in hidden services, which is actually in service.