A tightly-coupled distributed hypercube file system
Robert J. Flynn, Haldun Hadimioglu · 1991
A file is considered in a distributed memory hypercube system. A tightly-coupled distributed hypercube file system is proposed. The degree of the hypercube associated with the file system can range from 0 to n, the degree of the hypercube of processors. In this system, the degree of the hypercube disk system is considered to be smaller than that of the processor hypercube but larger than 0. Simulation of the resulting physically loosely coupled but from the perspective of the operating system tightly coupled distributed file system is applied to scientific applications. A distributed file system is presented and modeled. In this system large volumes are distributed to several disks, interconnected to each other and at the same time with specific disks more tightly coupled to clusters of processors of the hypercube. This raises the question of the interconnection of the disks and the organization of the files. Clearly this must be done in relation to applications, for it is the flow of data files at the request of the application's needs that must determine the appropriate number of processors for each disk and also the file flow among the interconnected disks. For selected examples, performance results are presented. By using the fact that typical hypercube applications have a crystalline structure, a distributed file system is formed where a set of trustworthy cooperative application processes interact with user-directed producer-driven file system processes. The user gives explicit instructions in application programs to the file system to move data between producer and consumer nodes which need to exchange little information prior to the data communication. In the distributed file system, increasing the number of disks brings more opportunities of overlapping of computations and file operations at the expense of increased data transmission. As more disks are added, the distributed file system performs better than the centralized file system due to the decrease in file access bottleneck. For larger numbers of disks, the performance degrades because of the increased demand for data on the host by the disks. This is reversed, however, when the disks move data more among themselves using the disks interconnection network.