Second Version of encT E X: UTF-8 Support
Petr Olÿ · 2003
The UTF-8 encoding keeps the standard ASCII characters unchanged and encodes the accented letters of our alphabets in two bytes. The standard 8-bit TEX is not ready for the UTF-8 input because it has to manage the single character as two tokens. It means you cannot set the \catcode, \uccode, etc., of these single characters and you cannot do \futurelet of the next character in the normal way. The second version of my encTEX solves these problems. The encTEX program is fully backward compatible with the original TEX. It adds ten new primitives by which you can set or read the conversion tables used by the input processor of TEX or used during output to the terminal, log and \write files. The second version creates the possibility of converting the multi-byte sequences to one byte or to a control sequence. You can implement up to 256 UTF-8 codes as one byte and an unlimited number of other UTF-8 codes as a control sequence. All internals in 8-bit TEX work as usual if the normal “one byte encoding” of input files is used. I think that the UTF-8 encoding will be used more commonly in the future. In such a situation, there is no other way than to modify the input processor of TEX; otherwise, the 8-bit TEX will be dead in a short time. Resume Le codage UTF-8 garde les caracteres ASCII inchanges et encode les lettres accentuees de nos alphabets en deux octets. Le TEX a 8 bits standard n’est pas pret pour une entree UTF-8 car il doit gerer les caracteres comme deux octets. Cela signifie que vous ne pouvez pas changer \catcode, \uccode, etc. de ces caracteres et vous ne pouvez pas faire un \futurelet du caractere qui suit, dans la sens habituel. La deuxieme version de mon encTEX resoud ces problemes. encTEX est totalement compatible avec le TEX original. Il ajoute dix nouvelle primitives par lesquelles vous pouvez etablir ou lire les tables de conversion utilisees par le processeur d’entree de TEX ou utilisees pendant la sortie au terminal, et aux fichiers log et \write. La seconde version donne la possibilite de convertir des sequences multi-octets vers un octet ou une commande. Vous pouvez implementer jusqu’a 256 codes UTF-8 comme un octet and un nombre illimite de codes UTF-8 comme commandes. Toute la machinerie interne de TEX fonctionne comme si les fichiers d’entree sont dans un «codage normal a un octet». L’auteur pense que le codage UTF-8 va etre de plus en plus courant dans l’avenir. Dans cette situation il n’y a pas d’autre moyen que de modifier le processeur d’entree de TEX, sinon la version originale de TEX a 8 bits va disparaitre sous peu.