MapReduce/Bigtable for Distributed Optimization

Keith Hall, Scott G. Gilpin, Gideon S. Mann · 2010

With large data sets, it can be time consuming to run gradient based optimization, for example to minimize the log-likelihood for maximum entropy models. Distributed methods are therefore appealing and a number of distributed gradient optimization strategies have been proposed including: distributed gradient, asynchronous updates, and iterative parameter mixtures. In this paper, we evaluate these various strategies with regards to their accuracy and speed over MapReduce/Bigtable and discuss techniques and configurations needed for high performance. 1

Read the paper · More papers on PaperTik