A statistical model for domain-independent text segmentation

Masao Utiyama, Hitoshi Isahara · 2001

We propose a statistical method that finds the maximum-probability segmentation of a given text. This method does not require training data because it estimates probabilities from the given text. Therefore, it can be applied to any text in any domain. An experiment showed that the method is more accurate than or at least as accurate as a state-of-the-art text segmentation system.

Read the paper · More papers on PaperTik