Toward a Representative DNS Data Corpus: A Longitudinal Comparison of Collection Methods

Calvin Kranig, Eric Pauley, Wei-Shiang Wung, Paul Barford, Mark E. Crovella, Joel Sommers · 2025

Domain Name System (DNS) records are frequently used to investigate a wide variety of Internet phenomena including topology, malicious activity, resource allocations and user behavior. The scope and utility of the results of these studies depends intrinsically on the representativeness of the DNS data. In this paper, we compare and contrast five different DNS datasets from four different providers collected over a period of 3 months that in total comprise over 10.1B total Fully Qualified Domain Names (FQDNs) of which 3.7B are unique. We process and organize that data into a consistent format that enables it to be efficiently analyzed in Google BigQuery. We begin by reporting the details of the measurement methods and the datasets used in our analysis. We then analyze the relative coverage of each dataset by structural, administrative, and client-behavioral features. We find that while there are significant overlaps in the records provided by each dataset, each also provides unique records not found in the other datasets that are important in different use cases. Our results highlight the opportunities in using these datasets in research and operations, and how combinations of datasets can provide broader and more diverse perspectives.

Read the paper · More papers on PaperTik