Distributed and Parallel Computing Issues in Data Warehousing (Invited Talk)
Héctor García-Molina, Wilburt Juan Labio, Janet L. Wiener, Y. Zhuge · 1998
A data warehouse is a repository of data that has been extracted and integrated from heterogeneous and autonomous distributed sources. The warehouse data is used for decision-support or data mining. In this paper we illustrate some of the challenges in distributed and parallel computing faced by such systems. Our examples come from research done in the Stanford WHIPS Project. 1 Introduction A data warehouse is a repository of data that has been extracted and integrated from heterogeneous and autonomous distributed sources. For example, a grocery store chain might integrate data from its inventory database, sales databases from different stores, and its marketing department's promotions records. The store chain could then: (1) find out how sales trends differ across regions of the country or world; (2) correlate its inventory with current sales and ensure that each store's inventory is replaced in keeping with its sales; (3) analyze which promotions lead to increased product sales. For...