Robust and Resilient Migration of Data Processing Systems to Public Hadoop Grid
Deepak Kumar Vasthimal · 2018
Data processing platforms ingest large and small datasets from various sources and process small datasets in real-time (<; 10 seconds) and large datasets in batch as defined by a job. Click stream data analysis involves collecting, analyzing and aggregating data for business analytics. A/B testing or any data experimentation uses click stream data stream to compute business lifts or capture user feedback to new changes on the site. These systems are hosted service and come with a framework to define and deploy some-kind of custom jobs to process data. Real-time processing takes place in-memory and batch processing takes place on distributed and scalable system like Hadoop. As Hadoop have gained popularity over the years to process large datasets raging from GBs to PBs there is ever increasing need to consolidate private Hadoop infrastructure spread across the company into a single, multi-tenant public Hadoop grid. In addition to cost effectiveness there can be multiple factors which can prompt the movement of data processing platforms from exclusive single tenant private grid(s) to a multi-tenant, large scale Hadoop grid. This can necessitate changes in data processing platforms but migrating existing jobs can pose multiple challenges. This paper discusses the motivation, challenges faced solutions employed and best practices. There is more focus on efforts taken to ease the migration of systems with least possible impact on customer.