Application-independent representation of text for document processing— —will Unicode suffice?
Frank G. Mittelbach, Chris Rowley · 1999
1 Abstract The large number, and the importance, of applications that support multi-lingual document processing have shown that the current methods and concepts for representing textual data in document processing systems are deficient in various respects. One major advance in this area is the development of the Unicode standard ([UC95]) for text character encoding. While this, or an similar updated standard, will play a prominent role in overcoming various problems in this area, contrary to common belief it is by no means all that is needed. This paper describes the requirements of an application-independent representation of textual data that provides a well-defined interface for applications and processes that act on such data. It shows what additions are needed to the Unicode standard for text character encoding [UC95] in order to provide a concrete syntax for such a representation.