N-gram Word prediction language models to identify the sequence of article blocks in English e-newspapers
Deepa Nagalavi, M. Hanumanthappa · 2016
In the analysis of newspaper page, an identification of individual article is an essential task. Since the articles are the most important information unit in a newspaper. A newspaper contains variety of multiple articles with different heterogeneous page layouts. Consequently the articles in a page are divided into multiple unordered blocks. In this paper the link is established between different blocks of an article with the reading order of a sentence. It is identified with an N-Gram based linguistic processing approach for the retrieval of individual article from newspaper. The model predicts the preceding word knowing the previous content with the probability of a word sequence.