Storage management in large distributed object-based storage systems

Scott Brandt, Feng Wang · 2006

Driven by the requirements for extremely high bandwidth and large capacity, storage subsystem architectures are undergoing fundamental changes. The object-based storage model, which repartitions the file system functionalities and offloads the storage management functions to intelligent storage devices, is a promising new model. In this model, Object-based Storage Devices (OSDs) manage their own storage space, provide persistent storage, and export an object interface to the data. With abundant on-board computing power, OSDs are able to provide complex services to facilitate high level system designs. This thesis focuses on the design, performance and functionality of an individual OSD in a large distributed object-based storage system, currently being developed in the Storage Systems Research Center at the University of California, Santa Cruz. Based on the file system workload analysis and the expected object workload studies, I extract unique features of object workloads and design an efficient storage manager, named OBFS, for individual OSDs. OBFS employs variable-sized blocks to optimize disk layouts and improve object throughput. Object metadata and attributes are also laid contiguously together with data, which further improves disk bandwidth utilization. The physical storage space is partitioned into fixed-size regions to organize blocks with different sizes together, which effectively reduces file system fragmentation in the long run. A series of experiments was conducted to evaluate OBFS performance compared with two Linux file systems, Ext2 and XFS. The results show that OBFS successfully limits the file system fragmentation even after long term aging. OBFS demonstrates very good synchronous write performance, exceeding those of Ext2 and XFS by up to 80% on both fresh systems and aged systems. Its asynchronous write performance is about 5% to 10% lower than that of Ext2, but 20% to 30% higher than that of XFS on a lightly used disk. On a heavily used disk, OBFS beats both Ext2 and XFS by 20%. For read operations, the performance of OBFS almost doubles that of Ext2 and is only slightly slower than that of XFS. Overall, OBFS achieves 30% to 40% performance improvements over Ext2 and XFS under expected object workloads, and forms a fundamental building block for larger distributed storage systems.

Read the paper · More papers on PaperTik