Improved document skew detection based on text line connected-component clustering
N. Liolios, Nikos Fakotakis, G. Kokkinakis · 2002
The classical method of document skew detection, based on nearest-neighbor clustering, is revisited. A heuristic is proposed which attempts to group all the connected components that belong to the same line of text, into one cluster. The larger clusters are known to result in better skew angle estimation. The skew detection accuracy of this improved connected-components method is several orders of magnitude better, when compared to the classical approach, with no change in the order of complexity.