Activity Transition Graph Generation: How Far Are We?
Jiakun Liu, Peixin Zhang, Han Hu, Yonghui Liu, Wei Minn, Ferdian Thung, Shahar Maoz, Eran Toch, Debin Gao, David Lo · ACM Transactions on Software Engineering and Methodology · 2025
Android applications (i.e., apps) are indispensable nowadays and are getting bigger and bigger with an increasing number of functionalities. To understand how to access functionalities in an app, prior studies proposed tools to model the transitions between functionalities with the activity transition graph (ATG). ATG is an important data structure and has been used for various Android app analyses, including app design, understanding, and testing. However, there is no benchmarking work on ATG generation. It is still unclear whether the transitions identified by tools are correct and how many transitions are missed. To fill this gap, we manually identified all transitions in 98 applications to build a benchmark. Using the benchmark, we evaluate seven popular ATG generation tools that were used in prior studies. We observe that these tools not only report incorrect transitions but also miss transitions, and different tools do not report the same set of transitions. Compared with the transitions reported by a single tool, the union set of the transitions reported by different tools contains fewer missed transitions but more incorrect transitions. We summarize five potential reasons that explain why GUI testing tools fail to identify transitions in ATG, revealing the limitations in the current design of exploration strategies. For instance, we observe that learning-based tools may overlook subtle distinctions between states, potentially misclassifying different states as identical. This will lead the tool to always focus on the old state while the new state is not fully explored. Based on our findings, we propose a series of suggestions for researchers using and building ATG generation tools. For example, when aiming to identify more transitions, one can combine the results of different tools by running each tool for 10 minutes, which will produce better results than running a single tool for 60 minutes.