Random Walks on the World Wide Web
Todd Silvestri · The Mathematica Journal · 2013
This article presents RandomWalkWeb, a package developed to perform random walks on the World Wide Web and to visualize the resulting data.Building upon the packageʼs functionality, we collected empirical network data consisting of 35,616 unique URLs (approximately 133,500 steps).An analysis was performed at the domain level and several properties of the web were measured.In particular, we estimated the power-law exponent g for the in-and out-degree distributions, and obtained values of 2.10 ± 0.09 and 2.36 ± 0.1, respectively.These values were found to be in good agreement with previously published results. ‡ 1 IntroductionThe World Wide Web (WWW), commonly referred to as simply "the web," is a vast information network accessible via the Internet.Initially proposed in 1989 [1], the web grew out of the work of Tim Berners-Lee while at the European Organization for Nuclear Research, known as CERN.Two software technologies form the core of the web, namely the HyperText Markup Language (HTML) and the Hypertext Transfer Protocol (HTTP).The HTML (or source) of a web page contains elements known as tags that describe the content of the document.For instance, the anchor tag defines a hyperlink, or link, to another document.Each file (or resource) on the web is identified by a Uniform Resource Locator (URL).A browser or other user agent may request a file via HTTP by specifying its URL.The response-typically the requested file-is again transmitted by HTTP from the web server to the client.The topology of large-scale complex networks, such as the web, can be explored using graph theoretic methods (see [2] and references therein).Specifically, the web can be viewed as a directed graph, where the web pages are vertices and the hyperlinks are edges.Unfortunately, two problems exist due to the nature of the web: (1) it cannot be indexed (or mapped) in its entirety; and (2) analyzing the corresponding graph would be highly computationally intensive.In fact, a recent announcement [3] suggests that the web may contain at least 10 12 unique URLs.