A new approach to sort Unicode Bengali text
Ahsanur Rahman, Md. Abdus Sattar · 2008
Character order in Unicode for Bengali is different from the sorting order suggested by the governing authority. As a result, simple letter by letter comparison doesn’t yield correct order of Bengali words. The presence of modifier characters in Bengali made the situation more complicated. The objective of our study is to adapt the suggested collation order for Unicode represented Bengali text while achieving maximum possible efficiency. Here we propose an algorithm for this purpose. The proposed algorithm is applicable to any chosen sorting order. Also it compares words in O(1) time, irrespective of their lengths. Thus complexity of sorting texts is always O(n log n).