A model and toolset for the uniform tagging of encoded documents
Julie A. Barnes, Sandra A. Mamrak · 1991
syntax of the LIF The construction of the abstract syntax of the LIF is driven by a partitioning of the token classes similar to that in Figure 3 for current encoding classes. The first partition of the token classes is that between tags and text. In a survey of current encoding schemes, one feature stands out as distinguishing tags and text. Each scheme reserves at least one printable keyboard character for the exclusive function of signaling the existence of a tag in the data stream. Scribe uses the symbol `@'. L A T E X [11] restricts the use of many characters, most notably `\\'. A reserved character ensures that the tag token classes and the text token classes are disjoint. We require that the LIF provide reserved characters to indicate the start and the end of a tag. Further partitions, as seen in Figure 4, eventually divide the tags into segmenting start tags, segmenting end tags, symbol tags, and nonsymbol tags. Because one of the goals in designing the LIF is that all marku...