Using Code Coverage to Assess Feature Gaps in MPI Correctness Tool Classification Tests
Alexander Hück, Simon Schwitanski, Tim Jammer, Joachim Jenke, Yussur Mustafa Oraji, Christian H Bischof · 2025
We examine the code generator-based MPI correctness benchmark MPI-BugBench (MBB) by analyzing the code coverage it triggers in three tools: MUST, PARCOACH, and clang-tidy. We present our analysis as a complement to MBB’s original design, which is based on pruning the potentially exhaustive test set based on real-world MPI usage patterns. Our assessment identifies two key limitations in MBB’s generated tests: incomplete coverage of MPI features, such as varying-count collectives, and limited structural diversity, such as lack of loops and array-based MPI handles. In addition, the current strategy of MBB’s code generation by increasing test volume alone offers limited benefit for exercising the analysis of these tools in our assessment. To address these gaps, we implemented 34 new tests covering missing MPI features and more varied code structures. The tests exercise previously uncovered analysis code, where a single varying-count collectives test adds 770 covered code lines in MUST.