Similarity Measures Between Strings Extended to Sets of Strings
Karen A. Lemone · IEEE Transactions on Pattern Analysis and Machine Intelligence · 1982
Similarity measures between strings (finite-length sequences of symbols) are extended to apply to sets of strings in an intuitive way which also preserves some of the desired properties of the initial similarity measure. Two quite different measures are used in the examples: the first for applications where numerical computation of pointwise similarity is needed; the second for applications which depend more heavily on substring similarity. It is presumed throughout that alphabets may be nondenumerable; in particular, the unit interval [0, 1] is used as the alphabet in the examples.