"All roads lead to Rome"

Sebastian Schelter, Stephan Ewen, Kostas Tzoumas, Volker Markl · 2013

Executing data-parallel iterative algorithms on large datasets is crucial for many advanced analytical applications in the fields of data mining and machine learning. Current systems for executing iterative tasks in large clusters typically achieve fault tolerance through rollback recovery. The principle behind this pessimistic approach is to periodically checkpoint the algorithm state. Upon failure, the system restores a consistent state from a previously written checkpoint and resumes execution from that point.

Read the paper · More papers on PaperTik