Data Exploration and Summaries
Emmanuel Paradis · 2020
Exploring data is an indispensable step before conducting model fitting or tests of hypotheses. It helps to assess the proportions of missing data, sample sizes per group, or the main patterns of variation in the data. Experience—and patience—are important during data exploration since some insights can be revealed by carefully looking at different facets of the data. Missing data can be a source of confusion, particularly with allelic data which can have several levels of complexity and where missing data can be coded in different ways. Pegas has a flexible framework with respect to missing data: most of the time, data are read “as is” in files. A useful approach to summarize information from a set of sequences is to first examine the sequences that are identical and calculate their frequencies. Pegas implements the class “haplotype” which works with different types of sequence data.