A Corpus-Based Approach to Text Partition
Kuang‐Hua Chen, Hsin‐Hsi Chen · 1995
A text partition model is proposed to determine the boundaries of discourse structures. It is based on association of noun-noun relations and noun-verb relations defined on discourse level and sentence level, respectively. Three factors are considered: 1) repetition of words, 2) importance of words, and 3) collocational semantics. A window is moved from the first sentence to the last one and the association norm for sentences in the current window is calculated. Finally, the peaks in the sentence position vs. association norm graph forms the potential discourse boundaries. Ten texts randomly selected from LOB corpus are used as the testing texts. The experimental results are compared with the readers ' judgment and the real boundaries in the testing texts. The applications of the results to sentence alignment, topic identification, topic shift and topic abstraction are discussed. 1.