Research on Automatic Abstracting of Chinese Web Page
Xu Yang Xiao · Computer and Modernization · 2006
Automatic abstracting is a practical and difficult branch in natural language processing,which becomes an important problem in domains such as Internet information retrieval.This paper describes an automatic abstract system to process Chinese Web page,which is mainly based on text structure.The method provided in this paper is to analyze the text structure firstly,obtain the positional information of the paragraph and all levels of subtitles information,then uses statistical methods and the heuristic rule to extract Key words and key sentences,and finally creates the abstract.Experiments show that this method can generate abstract effectively and efficiently.