Estimation of a Probability Distribution over a Hierarchical Classification

David McCarthy · 1997

Many natural language processing applications make use of hierarchical classifications whilst also having a statistical framework which requires an estimation of the probability distribution over the taxonomy. Data for estimating the probability distributions typically comes from corpora but estimation is complicated by the ambiguity of the data. One application involving such a task is the automatic acquisition of selectional preferences. The method of estimating these class probabilities is crucial to the success of the technique and this paper compares the previous schemes and concludes that correct modelling of class inclusion is essential. Modifications to previous approaches are described which help to curb the effect of ambiguity and avoid overly-general results.

Read the paper · More papers on PaperTik