On the effectiveness of the skew divergence for statistical language analysis

Lillian Lee · 2001

Estimating word co-occurrence probabili-ties is a problem underlying many appli-cations in statistical natural language pro-cessing. Distance-weighted (or similarity-weighted) averaging has been shown to be a promising approach to the analysis of novel co-occurrences. Many measures of distri-butional similarity have been proposed for use in the distance-weighted averaging frame-work; here, we empirically study their stabil-ity properties, nding that similarity-based estimation appears to make more ecient use of more reliable portions of the training data. We also investigate properties of the skew di-vergence, a weighted version of the Kullback-Leibler (KL) divergence; our results indicate that the skew divergence yields better results than the KL divergence even when the KL divergence is applied to more sophisticated probability estimates. 1

Read the paper · More papers on PaperTik