Scaling security for big, parallel file systems

Andrew W. Leung, Ethan L. Miller · File and Storage Technologies · 2007

The need for petaand exabyte scale parallel file systems that support high-performance computing (HPC) has been rapidly increasing. These systems have unique demands, different from those of traditional distributed file systems. As a result, securing I/O in big, parallel file systems without significantly impacting performance has proven challenging. Parallel file systems are commonly composed of three components: clients, metadata servers, and network-attached storage devices, such as disks or object-storage devices. Security is commonly based on a capability model, where capabilities are ’tokens’ which represent single block or object I/O authorization. Capabilities are generated by metadata servers, given to clients, and presented to storage devices with I/O requests. Petaand exabyte scale systems may have tens of thousands of clients and storage devices. Files are very large, often gigabytes to terabytes, and are commonly striped across thousands of storage devices. HPC workloads are parallel and bursty, meaning I/O often comes from many clients at a time with very short inter-arrival times. This implies files are commonly accessed by thousands of clients within a few seconds. These factors are often worst-case scenarios for many existing solutions. While there are many security schemes for distributed file systems, few have addressed large-scale, demanding environments. In these environments existing solutions cannot sustain high performance or must weaken security to do so. HPC workloads can burden metadata and data servers with generating and verifying millions of single block or object capabilities. Reliance on shared-key cryptography introduces vulnerabilities when millions of storage devices are susceptible to attack. Also, many common security techniques, such as revocation, become difficult with so many clients and storage devices.

Read the paper · More papers on PaperTik