An algorithm for conjunct identification in a natural language processing system

Rajeev Agarwal · 1992

This paper presents an approach taken by the author to identify the parts of the sentence that get conjoined by the co-ordinate conjunctions in English sentences. This identification is important for a natural language processing system for it to properly understand the sentence. The test domain being used for the algorithm presented here is a 10,000 word chapter of the Merck Veterinary Manual, which contains over 700,000 words. The primary features of the algorithm are that it is domain independent in nature and that it is being tested on such a large real-life domain.

Read the paper · More papers on PaperTik