Design and implementation of PDF reader
Shijin Liu · Jisuanji gongcheng yu sheji · 2010
To extract the text,images and graphical information from PDF file validly,an implementation model including four units(file pretreatment,display pretreatment,function extension and display) is raised.Based on the structure of PDF file,a solution of ignoring secondary message and positioning key information is put forward.On this basis,a solution to the data stream processed by FlateDecode,DCTDecode and CCITTFaxDecode filters is presented.After analyzed PDF pages twice,corresponding data structure of text and graphical are designed to record the results.At last the data utilization and function extension are discussed.The model can implement the extraction and display of information in PDF file well by experimental comparison,and it will benefit the further deve-lopment of PDF in the field of Chinese information processing.