Rio: a universal multimedia storage system based on random data allocation and block replication

José Renato Santos · 1998

We describe the design, implementation and performance analysis of the RIO (Randomized I/O) Multimedia Storage System which manages a set of parallel disks and supports real-time data retrieval with statistical delay guarantees. Previous work on multimedia storage systems has concentrated on video playback. RIO, however, is designed to support much more general workloads, such as 3D interactive virtual worlds, interactive scientific visualizations, etc., in addition to video and audio play-out. In RIO, data objects are divided into constant size data blocks, and data blocks are placed on randomly selected disks. This converts all logical access patterns to a uniform random physical access pattern to the parallel disks, and thus facilitates simultaneous support of multiple types of applications. To provide short term load balancing, RIO replicates an arbitrary fraction of data blocks on a different disk, also randomly selected. This data redundancy gives the disk scheduler flexibility in scheduling read requests, allowing the system to reduce the short term load imbalance, and achieve high throughput with low delay bound guarantees. We analyze the performance of RIO through simulations and experiments in our prototype for different systems parameters. Our results show that RIO has cost/performance competitive to that of striping techniques which is a traditional approach used in previous multimedia storage systems, with the advantage of supporting much more general workloads than striping techniques. Moreover, RIO outperforms striping techniques if some replication is used, even if replication is only partial. We discuss various schemes to support fault tolerance, and compare them in terms of reliability and real-time performance under normal operation and disk failure. We also discuss the performance of RIO under heterogeneous disk configurations and show that RIO provides a simple and effective way of dealing with heterogeneity. We also describe issues involved in extending RIO to a distributed environment, composed of a cluster of PCs. In a distributed environment, load balancing is complicated by the lack of instantaneous global knowledge of the disk queue lengths. We describe several schemes for load balancing with imperfect disk load knowledge and compare them in terms of scalability and real-time performance.

Read the paper · More papers on PaperTik