Do Query Optimizers Need to be SSD-aware?
Steven Pelley, Kristen LeFevre, Thomas F. Wenisch · 2011
Flash-based solid state disks (SSDs) are beginning to supplant conventional rotating disks for performance-critical data in myriad DBMS applications, including decision support systems. Though SSDs provide the same block-oriented storage abstraction as conventional disks, their performance characteristics differ drastically. Whereas SSDs provide relatively modest improvements in sequential transfer rates (e.g., perhaps 2 × improvement), they can provide over 100 × improvement for random reads, resulting in similar sustained transfer rates regardless of the access pattern. Conventional query optimizers assume a storage cost model where sequential IOs are far less costly than random IOs, and select access paths and join algorithms based on this assumption. Given the drastic change in SSD performance characteristics, intuition suggests that optimizer cost models must be updated (e.g., to prefer non-clustered index scans more frequently). Surprisingly, our empirical investigation using a commercial DBMS finds it is not necessary to adjust query optimization when shifting relations from disk to flash—an SSD-oblivious optimizer generally makes effective choices. We make two main observations. First, we demonstrate both empirically and analytically that the range of selectivities for which an unclustered index scan can benefit from SSDs ’ fast random reads is so narrow that it is inconsequential in practice. Second, our measurements show that the performance variations across alternative join algorithms on SSDs are generally smaller than the corresponding variation on disks and are dwarfed by the 5 × to 6 × performance boost of shifting data from disk to SSD. We conclude that existing query optimizers largely make correct decisions even when treating all storage devices as conventional disks, and the small and inconsistent performance gains available by making query optimizers SSD-aware are not worth the effort. 1.