Detecting Subject Boundaries Within Text: A Language Independent Statistical Approach
Korin Richmond, A. J. Smith, Einat Amitay · ERA · 1997
We describe here an algorithm for detecting subject boundaries within text based on a statistical lexical similarity measure. Hearst has already tackled this problem with good results (Hearst, 1994). One of her main assumptions is that a change in subject is accompanied by a change in vocabulary. Using this assumption, but by introducing a new measure of word significance, we have been able to build a robust and reliable algorithm which exhibits improved accuracy without sacrificing language independency. 1 Introduction Automatic detection of subject divisions within a text is considered to be a very difficult task even for humans, let alone machines. But such subject divisions are used in more complex tasks in text processing such as text summarisation. An automatic method for marking subject boundaries is highly desirable. Hearst (Hearst, 1994) addresses this problem by applying a statistical method for detecting subjects within text. Hearst describes an algorithm for what she calls...