Validating Analysis Data Set without Double Programming - An Alternative Way to Validate the Analysis Data Set
Linfeng Xu, Christina Scienski · 2014
This paper will demonstrate an alternative way to validate the analysis data set without double programming. The common practice for the most critical level validation of an analysis data set in pharmaceutical industry is double programming, that is, the source programmer generates an analysis data set (i.e. production data set) and another programmer (Reviewer) uses same specifications to create a QC data set. Reviewer then conducts PROC COMPARE to see whether or not there is any discrepancy between the two data sets. This paper will introduce the ALTVAL macro as an alternative to double programming in validating an analysis data set. The ALTVAL macro consists of two parts: 1) compare common variables between raw and analysis data set and use PROC COMPARE to check for discrepancies. (common variable check) 2) using unique merging variable(s), merge the raw and analysis data set to create a new data set, conduct a cross-frequency check between input variables and derived variables (cross check); 3) Based upon the step two created big combined data set, reviewer can write simple codes to report the cases which could not meet the variable derivation rules in the specification since we have input variables and derived variables in same big combined data set (the logic check of variable specification). This paper will discuss the conceptual design, macro parameters, and prerequisites using this new approach. It will also discuss what types of analysis data sets are suitable for this new validation approach. BACKGROUND Based on the importance of a data set, complexity of derived variables in an analysis data set, availability of resource, including manpower and time to conduct the validation, we usually assign varying levels of validation. For the most critical validation, double programming is the current standard practice for verifying whether or not the derived analysis data set is accurate and correct based on specification. A reviewer may also conduct code review, log check, and data set check in addition to the double programming. Double programming is very time-consuming and thus a very expensive way for validation. This paper describes an alternative approach in validating an analysis data set for the most critical level of validation, which can potentially save time and money substantially. One caveat is that this alternative approach will not replace the steps that involve code review, log check, and data set check.