Analyzing edits to static variability

Bittner, Paul Maximilian · OPen Access Repositorium der Universität Ulm (OPARU) (Ulm University) · 2025

Despite decades of successful software product-line engineering as a research discipline to study the development of highly configurable software, central operations are still missing a semantic foundation. One such operation is to generate individual variants of a software automatically from user-specified configurations. A famous implementation of this operation is the C preprocessor, which is used in the Linux kernel and many other open-source projects to adapt source code before compilation, i.e., statically. While researchers have proposed dozens of formal languages to model and study automated variant generation, there is no overview on which languages exist, how these competing notions of variability relate, and whether they share the same denotation. Crucially, it is unknown whether a shared semantic model of variability exists and if so, what properties it may have, analogous to how the lambda calculus or Turing machines are semantic models for programming languages. Similarly, despite being a central operation in developer workflows and despite being studied in a myriad of specific research problems, changes to variational software at the granularity level of patches and commits have no common semantic foundation as well. More so, there is neither a language nor a diff tool for expressing changes in a way that distinguishes edits to variability from edits to source code, and that is complete and generic (i.e., we may express all changes made to any language). Symptomatically, there is also a lack of generic and fine-grained variability-aware change impact analyses to assist developers in their daily workflows, for instance upon creating a commit or when reviewing a pull request. This thesis explores and identifies semantic foundations for static variability and edits to it, and based on that, develops three change impact analyses. First, we map the unknown landscape of formal variability languages. To this end, we distill an overview of variability languages from the literature and present a unified formalization of their syntax and semantics. By developing a meta-theory for semantics of variability languages, we explore their relationships and excavate a hierarchy of expressiveness. We find multiple existing variability languages to be semantically equivalent, and we find different levels of expressiveness. In particular, the choice calculus and its dialects are among the most expressive languages, substantiating its envisioned role as a lambda calculus for variability. Second, we formalize and implement variability-aware differencing to express fine-grained changes to variational software or, more precisely, changes to an expression in a variability language. In a variability-aware diff, a change's impact to variability information can be inspected and distinguished from changes to source code. We present strategies to turn any generic diff tool into a variability-aware differencer, such that existing tool support can be reused with little to moderate effort. With our theory and implementation, we provide a foundation for studying changes to variability in existing projects, developing change impact analyses, or making analyses on static variability incremental. Third, to put variability-aware differencing to practice and thereby validate our theory, we develop one proactive, one reactive, and one interactive change impact analysis. All three analyses are designed to be used by developers or by tools in developer's environments or workflows, to, for example, identify a change's effect to a particular variant, feature, or line of code. We implement all analyses and evaluate two of them on 1.7 million commits from the version histories of 44 open-source systems, including the Linux kernel, and find our analyses to provide almost instant feedback on real-world commits.

Read the paper · More papers on PaperTik