Supernetworking the metacomputer: enabling guaranteed bandwidth through deterministic and efficient provisioning
Silvia M. Figueira, Sumit Naiksatam · 2006
How does a protein fold? What happens to space-time when two black holes collide? What impact does species' gene flow have on an ecological community? What are the key factors that drive climate change? Did one of the trillions of collisions at the Large Hadron Collider produce a Higgs boson, the dark matter particle, or a black hole? Can we create an individualized model of each human being for targeted healthcare delivery? How do major technological changes affect human behavior and structure complex social relationships? What answers will we find---to questions we have yet to ask---in the very large datasets that are being produced by telescopes, sensor networks, and other experimental facilities? These questions---and many others---are only now becoming answerable because of advances in computing and related information technology 1. Once used by a handful of elite researchers in a few research communities on select problems, advanced computing has become essential for future progress across the frontiers of science and engineering. Powered by continuing improvements in microprocessor speeds, visualization, data systems, and collaboration platforms are poised to change the way research and education are accomplished. When all the resources needed for a computation experiment are not available at one place, sophisticated software technology has made it possible to aggregate these resources from geographically distributed locations into a metacomputer. The 'single location, single time zone' bottlenecks that plagued these valuable resources can now be eliminated. Applications like global scale weather simulations, nuclear fusion/fission simulations, genome analysis, or structural analysis and synthesis of proteins, which used to traditionally run only on supercomputers, can now be deployed on affordable commodity clusters. Some of these applications are not only computationally intensive, but are also data-intensive, requiring the daily exchange of Terabytes of data, and projected to reach Petabytes in the very near future. This data tsunami, i.e., the flood of data from high-performance computing (HPC) systems, has created an unprecedented challenge for the data communication and networking infrastructure. The success of the Internet has greatly surpassed the expectations of its creators, but it is simply not suited to handle this deluge of data. A radical new approach is sought, which should not only meet the colossal requirements of data-hungry applications, but also serves to expose the network as a pliable resource. The later requirement, especially, is a critical one to address the paradigm shift to service-oriented computing. This ubiquitous supernetwork then replaces the computer as the heart of a new digital universe of billions of distributed computational elements and storage devices. The notion of LambdaGrids, powered by high-speed dynamic optical networks, is fast emerging as an exciting solution to these networking requirements. Technological advances in the field of photonics and management software now make it possible to orchestrate the tremendous bandwidth potential of optical networks with great deal of finesse. This dissertation explores how these new advances can be exploited for the realization of LambdaGrids. We aim to address several system-wide issues in achieving this objective and follow a holistic approach, integrating architecture, design, implementation, optimization, tools, and usability. The network footprint of HPC applications show pronounced peaks and valleys in utilization, prompting an overhaul of the traditional network provisioning styles such as peak-provisioning, point-and-click, and operator-assisted provisioning. A service-oriented stack must become capable of dynamically orchestrating a complex set of variables related to application requirements, data services, and network provisioning services, all withina rapidly and continually changing environment. Presented in this dissertation is DWDM-RAM, a prototypical platform implemented on a unique wide-area optical networking testbed incorporating state-of-the-art photonic components. We present results of research conducted on this prototype, which indicate that these methods have the potential to address several major challenges related to data-intensive applications. The central resource in the DWDM-RAM architecture is dynamically-provisioned bandwidth. Data channels riding on this bandwidth can be reserved in advance and are guaranteed to be available at the required time. This scheme of advance reservations is novel, and LambdaGrid-based applications, which can benefit from it are only now emerging. We formally define and analyze this scheme and present a constrained mathematical model for advance reservations. We also introduce FONTS - the Flexible Optical Network Traffic Simulator, a tool for simulating advance reservation, on-demand, and periodic data transfer requests. FONTS is based on a stochastic model and incorporates a variety of variables, which have been identified to accurately model advance reservation requests. FONTS validates the mathematical model and also helps to analyze complex scenarios. FONTS is available online and has been enthusiastically received by the research community. It has been shown in the context of the Internet that the distribution of sizes of transferred files significantly impacts the behavior of the underlying network and its throughput. Architects of emerging LambdaGrid-like environments also have to take into consideration the variation in file sizes when modeling the network infrastructure. They need to test, by simulation, resource management strategies which can efficiently transport and store these files. Hence, there is a need to model and generate representative workloads of files which will serve as input for these simulations. In this dissertation, we present a chaos-theoretic approach involving single intermittency maps for generating file sizes, which follow the generalized Zipf's law. Our objective is to present an alternative to the commonly used inverse transformation method for generating generalized Zipfian distributions. To experiment with intermittency maps, we have developed an online tool with which we observe that the chaos-theoretic approach leads to a larger dynamic range of file sizes as compared to the inverse transformation method. To mitigate the problem of bandwidth fragmentation in LambdaGrids we introduce the concept of elastic reservation of bandwidth capacity and present a network model which can support the elastic reservations. We also define the Elastic Scheduling Problem (ESP), which succinctly captures the optimal utilization objective of elastic reservations. Analysis of ESP reveals that it is an NP-complete problem. Hence we present a heuristic algorithm, Squeeze In Stretch Out (SISO), for tackling ESP. SISO achieves good bandwidth utilization in simulation and handles efficiently the dynamic sharing of bandwidth between advance and immediate reservation requests. In general, the approach for elastic reservation and retrospective scheduling presented in this dissertation is applicable to any shared resource with quasi-flexible characteristics. The elastic reservations model makes an implicit assumption about the quasi-flexibility of the user requests, but at times this quasi-flexibility may have to be induced with appropriate incentives. We explore how the Network Service Provider (NSP) can encourage user flexibility by dynamically engineering pricing incentives and by suggesting request realignment to overcome reservation contention. Ultimately, user flexibility leads to efficient network utilization, reduces the price for the users, and increases the revenue for the NSP. As history has shown, the scientific community has often led the way for technological advances that in the end become commonplace for the general public; the birth of the Internet being a perfect example of such evolution in the past. We believe that work, such as the one presented in this dissertation, will provide the necessary insight to the research community, and help in accelerating the endeavor of bringing high-performance computing within the reach of the masses. This is the over-arching sentiment of this dissertation. 1NSF'S 'CYBERINFRASTRUCTURE VISION FOR 21 st CENTURY DISCOVERY NSF', Cyberinfrastructure Council National Science Foundation (http://www.nsfgov/od/oci/CI-v40.pdf).