Efficient Large-Scale Distributed Training of Conditional Maximum Entropy Models

Ryan McDonald, Mehryar Mohri, Nathan Silberman, Dan L. Walker, Gideon S. Mann · Neural Information Processing Systems · 2009

Training conditional maximum entropy models on massive data sets requires significant computational resources. We examine three common distributed training methods for conditional maxent: a distributed gradient computation method, a majority vote method, and a mixture weight method. We analyze and compare the CPU and network time complexity of each of these methods and present a theoretical analysis of conditional maxent models, including a study of the convergence of the mixture weight method, the most resource-efficient technique. We also report the results of large-scale experiments comparing these three methods which demonstrate the benefits of the mixture weight method: this method consumes less resources, while achieving a performance comparable to that of standard approaches.

Read the paper · More papers on PaperTik