Making Legacy Digital Content Accessible at Source
Sankalan Pal Chowdhary, Dipendra Manocha, S. Balakrishnan, Akashdeep Bansal, Himanshu Garg · 2019
Nearly three decades have passed since the Unicode standard was first published in 1991. A lot of electronically generated content is still locked inside legacy encodings. This can be attributed to lack of software support for Unicode at the time of content creation. The target output originally was print and this went on unnoticed. Later, to meet the growing demand for digital content, the same content in legacy encodings had to be exported to unsearchable PDF's/EPUB's. Conversion to Unicode has been a challenge because digital publishing applications cannot provide built in conversion support for the multitude of legacy encodings. Conversion tools, even where available, are external to the source application and require manual effort not only for text export/import but also for correcting errors in conversion and changes in document layout.