Architecture of the LHCb Distributed Computing System

Philippe CHARPENTIER, Federico Stagni · 2016

The LHCb Distributed Computing system is based on the DIRAC interware, and includes many LHCb-specific extensions to form the BeautyDIRAC package.BeautyDIRAC is implementing the LHCb computing model.It handles workflows for all the distributed computing activities of LHCb.BeautyDIRAC provides extensions of the DIRAC components and interfaces, including a secure web client, python APIs and CLIs.The LHCb Production Management System (PMS) exposes to physics teams a powerful interface allowing the submission of complex requests for tasks involving a large number of jobs (productions).The PMS manages all types of LHCb production activities: simulation, reconstruction of real and simulated data, physics selection of events (stripping), working group analysis and event indexing.Automatized testing phases are implemented for simulation productions, as well as productions' validation and completion.Managers of simulation productions can therefore test, submit and verify a large number of production requests with a minimal effort.The testing phase is both a functional and a performance test, using a dedicated testing facility.Physics groups requesting productions can follow their progress through a user-friendly web interface.The PMS is built on top of the LHCb extensions of the Dirac Data Management (DMS) and Workload Management systems (WMS), which are highly integrated through the DIRAC components.The productionsâ Ȃ Ź tasks are handled by the BeautyDirac extension of the Dirac Transformation System (TS) from their creation to their completion.Tasks requiring input datasets are fully data-driven (using the DMS), as new files becoming part of a dataset can automatically generate the creation of new tasks.It is therefore quite easy to submit chains of productions, some of which consume as input dataset the output datasets of others.For simulation production chains, an agent is also in charge of creating new tasks until the required number of simulated events is available.The tasks are submitted as jobs to the WMS and eventually run on a large palette of resources (Grid sites, Cloud sites, HPC centers, computer clusters, volunteer computing platformsâ Ȃę).User jobs are submitted through the same WMS as production jobs, which allows prioritizing jobs and running seamlessly user and production jobs within the same "pilot jobs", that constitutes a resource overlay in the DIRAC WMS.In this contribution we shall describe the synergy between the various BeautyDIRAC components (DMS, TS, WMS, PMS).We shall present in some details the LHCb Production Management System and show how it is used by a large community of physicists.We shall give examples on how very large productions can be centrally handled very efficiently by a small operations team.Finally, we shall give an overview of the LHCb Computing Operations' successes of the past few years.

Read the paper · More papers on PaperTik