String Comparison in XSLT with tan:diff()
Joel Kalvesmaki · Balisage series on markup technologies · 2021
Classical models of string comparison have been difficult to implement in XSLT, in part because those models are designed for imperative, stateful programming. In this article I introduce tan:diff() , an XSLT function built upon a different approach to string comparison, one more conducive to a declarative, stateless language. tan:diff() is efficient and fast, even on pairs of very long strings (100K to 1M characters), in part because of its staggered-sample approach, in part because of its strategies for optimizing enormous strings (> 1M characters). The output is qualitatively excellent: the function normally returns a minimal diff (shortest edit script) and the longest common substring. Released as part of an open-source function library, tan:diff() enables developers to incorporate robust text comparison directly into XML applications.