Data capture & conversion

Mark Gross · Columbia University Press eBooks · 2003

This chapter discusses the ins and outs of converting and upgrading textual data and related information from all kinds of formats into markup languages for publishing purposes. While the issues and approaches apply to all markup languages, particular emphasis is given to the issues of converting documents into XML, SGML, and HTML. The chapter looks at the issues surrounding the capture of data from sources such as paper and microfilm, word processors, publishing systems, PDF, and ASCII--the main issue being the need to untangle content from structure.

Read the paper · More papers on PaperTik