The Security of Latent Dirichlet Allocation
Shike Mei, Xiaojin Zhu · 2015
Latent Dirichlet allocation (LDA) is an in-creasingly popular tool for data analysis in many domains. If LDA output affects de-cision making (especially when money is in-volved), there is an incentive for attackers to compromise it. We ask the question: how can an attacker minimally poison the corpus so that LDA produces topics that the attacker wants the LDA user to see? Answering this question is important to characterize such at-tacks, and to develop defenses in the future. We give a novel bilevel optimization formu-lation to identify the optimal poisoning at-tack. We present an efficient solution (up to local optima) using descent method and im-plicit functions. We demonstrate poisoning attacks on LDA with extensive experiments, and discuss possible defenses. 1