Transformation of structured documents
Eila Kuikka, Martti Penttonen · 1996
Structure de nitions of documents have been used successfully for inputting and formatting in text processing systems. This report considers transformations between different representations of structured documents and studies possibilities to extend the use of structure de nitions to document transformations and to discover algorithmic methods for carrying out transformations. Documents are presented as parse trees for context-free grammars and transformations are made from parse tree to parse tree. First, the report describes di erences of manuscript styles required by various scienti c journals and presents a declarative classi cation for structure di erences between two parse trees. Second, a set of tree transformation methods are described and their suitability for transformations between documents having a structure di erence in each de ned class is analyzed. For each class several methods may ormust be used and only certain kinds of di erences can be managed automatically. Finally, instead of designing a system where a method accommodates for all kinds of di erences or where di erent methods are used in various transformations, the report presents a model for a document transformation system that presents a possibilityof using various methods according to di erences in document representations. The system is divided two modules. In the rst one transformations are made automatically and they do not change the hierarchical structure of a document. In the second one transformations are made semiautomatically or nonautomatically and the hierarchical structure changes. Differences between the existing and the required representation of a document are analyzed and methods selected according to the classi ed di erences.