Designing Graphics Requires Useful Experimental Testing Frameworks and Graphics Derived from Empirical Results
Susan VanderPlas · Harvard Data Science Review · 2021
From Empirical Results 2Hullman and Gelman (2021, this issue) have provided a very thorough discussion of the premise that interactive exploratory data analysis requires a theoretical framework for graphical inference to effectively support the analyst and counter any tendencies towards assuming all results are real and not just due to sample variability.They specifically mention a common problem with papers that theorize about interactive analysis:Even if some activities fall outside of the predictions of any specific model, without an underlying theoretical framework to guide the design of tools, we are hard pressed to identify where our expectations have been proven wrong and can easily end up with the sort of piecemeal and mostly conceptual theories that dominate much of the literature on interactive analysis.This lack of formalization makes it difficult to falsify or derive clear design implications from theoretical work.Unfortunately, the theoretical framework for model checking during exploratory and confirmatory data analysis proposed in this paper is just another conceptual and theoretical framework that is difficult to test or falsify as presented.The authors clearly support empirical validation of graphical analysis, but not enough to empirically examine their proposals via testing even a simple mock-up implementation of such a system with a fixed, relatively simple, data set.Without this empirical analysis, it is very difficult to see what this proposal adds to the two empirical methods discussed within as sub-cases of the model-check system, Bayesian Cognition and Visual Inference.These two formulations of empirical graphical testing arise out of different goals-to understand the mathematical reasoning used when processing charts, and to assess the perception of charts in the presence of randomization based on a null model.Unfortunately, the model-check integration proposal mentions, but does not actually address or even propose concrete solutions, to several major challenges that would need to be overcome in order to integrate either option into an automatic system for on-the-fly model checking.Of the two empirical processes 'subsumed' into the Bayesian model-check formulation, visual inference is perhaps the closest to the goals of model-checking integration proposed in this paper; as this is where I have done the majority of my work, I will address the shortcomings of the proposed system primarily from this perspective.