Web Page Design and Download Time.
Jing Zhi · Int. CMG Conference · 2001
Many factors contribute to Web site performance, most of which are at least partially outside the control of the site designer. Web page download times depend on page design, on Web server and client hardware and software configurations, and on the performance characteristics of the Internet route connecting a client to the site [Neil2000]. Of these, only page design is truly under the site designer’s control. However, if we assume a user with a high-speed connection and a site whose servers are not overloaded — as is typical of a business to business (“B2B”) interaction — then only two significant factors remain: site design and Internet latency between the client and the server [Sper1995] [Heid1997] [Touc1998]. In this paper we analyze measurement data based on test pages to explore various relationships between Web page design and page download time. Concentrating on information about the page and measures of Internet round trip time, we develop several specialized formulae to predict typical page download times in a B2B environment. After some introductory discussion of Web download components and experimental setup, the first part of this paper identifies packet count, rather than page size, as the crucial predictor of download time. We indicate how to calculate packet count in the absence of packet sniffer software. Next, we explore page download time as a function of page size, in a single-threaded environment. Here we build two different linear models to understand the bulk of page download for simple test pages (“Experiment A”). Then, we investigate how multi-threading improves performance for more complex pages. It is particularly challenging to build accurate models for multi-threaded behavior since the distribution of download activity among threads can be altered substantially by a single Web page element that loads unusually slowly. A more complicated model is built and discussed for pages with one to 64 embedded images (“Experiment B”). Since Internet performance varies erratically over time even for a fixed client-server pair, we discuss tradeoffs that can be made in modeling performance based on all available measurements vs. a more well-behaved subset containing approximately 90% of available measurements. We find that these experimental results represent and explain the basic mechanisms that regulate many of the performance characteristics observed while working with the extensive set of measurements collected annually by the author’s company.