Geographically distributed Batch System as a Service: the INDIGO-DataCloud approach exploiting HTCondor
Cristina Aiftimiei, Marica Antonacci, S. Bagnasco, T. Boccali, R. Bucchi, Miguel Caballer, Alessandro Costantini, Giacinto Donvito, Luciano Gaido, Alessandro Italiano, Diego Michelotto, M. Panella, Davide Salomoni, Sara Vallero · Journal of Physics Conference Series · 2017
One of the challenges a scientific computing center has to face is to keep delivering well consolidated computational frameworks (i.e. the batch computing farm), while conforming to modern computing paradigms. The aim is to ease system administration at all levels (from hardware to applications) and to provide a smooth end-user experience. Within the INDIGO- DataCloud project, we adopt two different approaches to implement a PaaS-level, on-demand Batch Farm Service based on HTCondor and Mesos. In the first approach, described in this paper, the various HTCondor daemons are packaged inside pre-configured Docker images and deployed as Long Running Services through Marathon, profiting from its health checks and failover capabilities. In the second approach, we are going to implement an ad-hoc HTCondor framework for Mesos. Container-to-container communication and isolation have been addressed exploring a solution based on overlay networks (based on the Calico Project). Finally, we have studied the possibility to deploy an HTCondor cluster that spans over different sites, exploiting the Condor Connection Broker component, that allows communication across a private network boundary or firewall as in case of multi-site deployments. In this paper, we are going to describe and motivate our implementation choices and to show the results of the first tests performed.