Notes on the KL-divergence retrieval formula and Dirichlet prior smoothing

ChengXiang Zhai · 2003

Given two probability mass functions p(x) and q(x), D(p jj q), the Kullback-Leibler divergence (or relative entropy) between p and q is defined as D(p jj q) = X x p(x) log p(x) q(x) It is easy to show that D(p jj q) is always non-negative and is zero if and only if p = q. Even though it is not a true distance between distributions (because it is not symmetric and does not satisfy the trian-gle inequality), it is still often useful to think of the KL-divergence as a “distance ” between distributions [Cover and Thomas, 1991]. 2 Using KL-divergence for retrieval Suppose that a query q is generated by a generative model p(q j Q) with Q denoting the parameters of the query unigram language model. Similarly, assume that a document d is generated by a generative model

Read the paper · More papers on PaperTik