Using Large Language Models to Support the Workflow of Differential Testing

Arun Krishna Vajjala, Ajay Krishna Vajjala, Carmen Badea, Christian Bird, Jade D'Souza, Robert A DeLine, Mikhail O Demyanyuk, Jason Entenmann, Nicole Forsgren, Aliaksandr Hramadski, Haris Mohammad, Sandeepan Sanyal, Oleg Surmachev, Thomas Zimmermann · 2025

Many software development teams use differential testing as a quality gate in their release process. Differential testing—namely, comparing behavioral differences between a system in production and a system in test—is a laborious process to label changes as regressions, expected changes, or incidental changes (e.g. those due to nondeterminism or timing). This manual process involves inspecting large textual artifacts, like logs, pull requests, and team discussions, which suggests that Large Language Models (LLMs) could be helpful. In this paper, we engage with the team developing a central Azure service to understand their work practice for differential testing. We used a design probe method to elicit feedback about several ways to use LLMs to improve their work practice, including automatically labeling behavior differences and providing summaries of various artifacts and discussions. Release engineers on the team report that predicting a difference's label would save them effort, but they want an explicit rationale to improve their trust in the prediction; they found the generated summaries to be informative, if a bit wordy.

Read the paper · More papers on PaperTik