Experience with the Parallel Workloads Archive
Dror G. Feitelson, Dan Tsafrir, David Krakov · 2012
Workload traces from real computer systems are invaluable for research purposes but regularly suffer from quality issues that might distort the results. As uncovering such issues can be difficult, researchers would benefit if, in addition to the data, the accumulated experience concerning its quality and possible corrections is also made available. We attempt to provide this information for the Parallel Workloads Archive, a repository of job-level usage data from large-scale parallel supercomputers, clusters, and grids, which has been used extensively in research on job scheduling strategies for parallel systems. Data quality problems encountered include missing data, inconsistent data, erroneous data, system configuration changes during the logging period, and unrepresentative user behavior. Some of these may be countered by filtering out the problematic data items. In other cases, being cognizant of the problems may affect the decision of which datasets to use.