Uncovering and exploiting the intrinsic correlations between file references
Dan Duchamp, Hui Lei · 1998
Distributed file systems provide a mechanism for physically dispersed users to share data and storage resources. Two recent technology trends have imposed new requirements on the performance and availability of distributed file service: the rapid improvements in processor and memory speeds, and the proliferation of mobile computers. The former creates an ever-widening performance gap between file I/O and the rest of the computer system. The latter makes voluntary disconnections from wired network environments a common phenomenon, during which file availability has to be optimized. A critical architectural feature of distributed tile systems is the caching of data at clients. It has been recognized that client-side caching can be exploited to address the performance and availability issues. In particular, prefetching increases cache hit rate and reduces file read latency; hoarding fills the client cache with useful data prior to a mobile client's disconnected operation so that the client can service most file accesses from its cache while disconnected. In this dissertation, we demonstrate that intrinsic correlations between file references exist and can be utilized in file prefetching and hoarding. We first develop a technique called semantic correlations, which monitors interesting system events and captures information that serves to structure and make sense of file references. Equipped with this technique, we then design, implement and evaluate a file prefetching mechanism and a file hoarding mechanism. Our prefetching mechanism is completely transparent to applications and users. It predicts future file references by recognizing directory traversals and by relating to past file usage patterns, based on an understanding of the file reference context. It reduces read latency by up to 35%, miss rate by up to 82% for the attribute cache and 46% for the buffer cache. Our hoarding mechanism intelligently detects files that are hidden from the user, and allows the user to gather and supply hoarding information from a high level. It provides high availability of files at small costs and without undue burden on the user.