An investigation of different string coding methods
Pankaj Goyal · Journal of the American Society for Information Science · 1984
Abstract This study Investigates various techniques for the automatic coding of English Language strings. The strings are titles drawn from bibliographic files, but the techniques do not require any prior knowledge of the source. Variations of the techniques have also been tested and it has been shown that string abbreviation or coding based on the maximum entropy principle gives the best results. Some of the reasons for the nonuniqueness of the generated codes are presented.