Ceph Parallel File System Evaluation Report

Feiyi Wang, USDOE Office of Science (SC), Mark Nelson, H. Oral, Douglas Fuller, Scott Atchley, Blake Caldwell, James A Simmons, Bradley W. Settlemyer, Jason Hill · 2013

We used Data Direct Networks’ (DDN) SFA10K as the storage backend during this evaluation. It consists of 200 SAS drives and 280 SATA drives, organized into various RAID levels by two active-active RAID controllers. The exported RAID groups by these controllers are driven by four hosts. Each host has two InfiniBand (IB) QDR connections to the storage backend. We used a single dualport Mellanox connectX IB card per host. By our calculation, this setup can saturate SFA10K’s maximum theoretical throughput (~12 GB/s). Our Ceph testbed employs a collection of testing nodes. These nodes and their roles are summarized in Table 1. In the following discussion, we use “servers”, “osd servers”, “server hosts” interchangeably. We will emphasize with “client” prefix when we want to distinguish it from above. All hosts (client and servers) were configured with Redhat 6.3 and kernel version 3.5.1 initially, and later upgraded to 3.9 (rhl-ceph image), Glibc 2.12 with syncfs support, locally patched. We used the Ceph 0.48 and 0.55 release in the initial tests, upgraded to 0.64 and then to 0.67RC for a final round of tests.

Read the paper · More papers on PaperTik