Sorting and measures of disorder
Vladimir Estivill‐Castro · 1992
It has been observed that, in practice, sequences to be sorted are often nearly sorted. Nearly sorted or presorted sequences are regarded as easier instances of the sorting problem and sorting algorithms that perform a number of comparisons that is a nondecreasing function of the size and the difficulty of a problem instance are called adaptive. This thesis is concerned with adaptive sorting algorithms. Existing order in a sequence is evaluated using what we call a measure of disorder. If a measure of disorder satisfies some additional properties we obtain what has been called a measure of presortedness. We explore the axiomatic definition of measures of presortedness and how it reflects intuition, and we design schemes to compare different measures of disorder. The relationship between right invariant metrics, used in Statistics to evaluate; correlation in ranking models, and measures of presortedness is studied. One result of this investigation is that we are able to obtain a general method to generate pseudo-random nearly sorted sequences. A new measure of presortedness related to sorting in mesh-connected processor arrays is defined and we present optimal sorting algorithms for this measure. We also give a generic adaptive sorting algorithm that is the basis of anew adaptive sorting algorithm that is not only practical, but also optimal with respect to five important measures of presortedness. We characterize, for the first time, the adaptive behavior of two variants of Quicksort. Motivated by these results, and since previous work used only worst-case analysis, we explore expected-case optimality with respect to measures of presortedness. We establish the corresponding theory and present three new sorting algorithms and their analyses: Randomized Quicksort is shown to be adaptive with respect to Exchange, in the expected case; Randomized Mergesort is shown to be Runs-optimal in the expected case; and Skip Sort is shown to be optimal with respect to five important measures in the expected case.