Determining file and directory importance

Kong-Cheng Wong · 1991

File replication is the most popular approach used to promote system reliability and file availability in a network based environment. However, all of the distributed file systems equipped with the functionality of file replication require their system users to determine how important their files are, in order to assist systems in making decisions on distributing replicas in the network (Blair et al., 1987). As such, system users are inevitably burdened with this potential responsibility. The problem can be partially alleviated if the system can take more responsibility for their system users in determining file importance. To achieve this goal, however, we need to better understand how system users cognitively make decisions regarding determining file importance. We first quantitatively compare the performance of three decision-making models popularly used in juror decision-making (Pennington and Hastie, 1981) to examine how satisfactorily they model the process of determining file importance. The three models are the linear weighting model, the Bayesian model, and the Poisson model. We then propose a simple, yet powerful, decision-making model, which is called the predictor domination model, for determining file importance. The model proposed suggests that the maximum predictor values observed in the session of determining file importance may be taken as the file importance. We next examine how significantly domain-dependent information contributes to determining file importance. We demonstrate using the linear weighting model that domain-dependent information seems to contribute non-negligibly to determining file importance. Since directories are usually treated as files used to store necessary information for other files, including directories, we therefore examine how directory importance can be determined. Since a file is locatable only through its corresponding pathname defined by its associated tree-structured directory system, the importance of a particular directory is determined by its child files and directories having the highest importance ratings. It is also suggested that grouping those files having a higher file importance near the root will save not only file access time, but also the space needed for storing directory structures.

Read the paper · More papers on PaperTik