Calculating the Posterior Odds from a Single-Match DNA Database Search under various Scenarios with Minimal Assumptions

Ronald W. J. Meester, Klaas Slooten · Law Probability and Risk · 2019

In their paper, “Rejoinder for Calculating the Posterior Odds from a Single-Match DNA Database Search” (Rouder et al., 2019), Rouder, Wixted and Christenfeld (RWC in the sequel) reply to our criticism (Meester and Slooten, 2019) on their paper (Wixted et al., 2019). They make the allegation that our reaction contains material that is, according to them, ‘simply wrong’. In this article, we reply to this allegation, showing that we were not wrong, and carefully explaining which erroneous suppositions in Wixted et al. (2019) led to their allegation. We also think that it is important to avoid that irrelevant notions (such as introduced in Wixted et al. (2019)) enter the discussion of how to deal with unique database matches. We will therefore sketch how to evaluate posterior odds against a suspect found in a database search in various scenarios, making as few assumptions as possible. We will see that populations or their sizes never come into play. In Meester and Slooten (2019), we criticized the way that Wixted et al. (2019) derive the posterior odds corresponding to a single match in a DNA database D⁠, not because of the obtained result, but because it is unnecessarily complicated. In their analysis, the size of the active criminal population plays a central role, whereas we claim that it does not, and that this fact removes the need for assumptions that the authors have made. We agree therefore with RWC that our comments ‘[…] strengthen our argument in favour of reporting case independent posterior odds’. Indeed, this is so, precisely because the most controversial estimate (the size of the active criminal population), which they considered to be necessary and which also led to a substantial part of criticism of the other commentators, is actually not an ingredient in the posterior odds. Before we go into some detail, we first recall that in our reaction, we have explicitly reproduced the calculations of Wixted et al. (2019) ‘without the size of the active criminal population’; we would expect that alone to be sufficiently convincing. If we do not need it, how can it play a role? Apparently this was not so, and RWC essentially claim that we indirectly estimate that size, since, as RWC claim, estimating the probability P(C∈D) (the probability that the database contains the trace donor) is tantamount to estimating the size of the active criminal population. This, however, is a misunderstanding. This is only so ‘within the framework that RWC themselves have developed’. As RWC write, they ‘needed to make some substantive assumptions, including that the profiles in the database are a random sample from some larger sample (which we call the active criminal population) and that each active criminal has the same probability of being in the database’. It is precisely this set of assumptions that is not needed, as we will motivate further in this commentary. No further assumptions have been made beyond assuming that all we know about S is that it is the only matching person. The equation above is equivalent to (2) in RWC, which they do not call into question. The point of debate therefore seems to be whether in this expression, it is necessary to interpret w as a measurement of the size of the active criminal population. Clearly, this is not so, since we derived it without this assumption. The quantity w is a proportion of previous samples that yielded a match. We plug in w as an estimate for P(C∈D)⁠, simply because we take as probability to obtain a match with the trace donor, the previous proportion of traces that gave a match; implicitly, assuming that this estimate works since (almost) all matches were with the trace donors and that the past is predictive for the future. RWC claim that in order to assess P(C∈D)⁠, ‘One answer is to make substantive assumptions, including that the profiles in the database are a random sample from some larger sample (which we call the active criminal population) and that each active criminal has the same probability of being in the database. Then, under these conditions, we can use w as an estimate of P(C∈D)⁠.’ However, while these assumptions are sufficient to interpret w as such, they are not needed at all, and we have nowhere made them to derive (4). As RWC write, this is one answer. It is not ours. For a database of a given size, it is immaterial for the posterior odds which fraction of the database comes from the active criminal population. When we have P(C∈D)⁠, it may be that only a small fraction in the database is from the active criminal population, or that a large fraction is, but whatever it is, P(C∈D) is what we need. Given P(C∈D)⁠, as long as we do not know who S is, the posterior odds are as above. The fraction of D that comes from the active criminal population is important to infer an estimate of the size of the active criminal population from w, but it is not for P(C∈D)⁠. Or, put differently still, if we want to phrase the required probability P(C∈D) in terms of the criminal population then what is relevant is the database ‘coverage’ of that criminal population, and not its size. Surely, making additional assumptions one can make inferences about the size from the coverage. The problem we have with Wixted et al. (2019) is not so much that they establish this link, but that they claim that the evaluation of the size of the criminal population is needed for the posterior odds. It is not, only the database coverage is. RWC claim that (at the end of section 1.1) ‘both approaches yield exactly the same numerical values because they rely on exactly the same assumptions. Importantly, in evaluation, they both rely on a concept of the size of some population, N.’ We now see that this is simply not true: we just showed how these assumptions are not necessary. RWC conclude the necessity of these assumptions from their sufficiency. We prefer to work without and use w directly for P(C∈D)⁠, all the more since with or without the assumptions of WCR, the posterior odds amount to (4). We do not object to phrasing w as a coverage of the criminal population, but also do not feel that such an interpretation is very helpful. Indeed, the concept of active criminal population is slippery. Introducing it brings in room for discussion which is, for the posterior odds of a database match, off topic. For example, Neumann and Ausdemore (2019) criticize the assumptions made in Wixted et al. (2019) regarding the active criminal population. That criticism has impact on the estimate that Wixted et al. (2019) propose but not on the posterior odds (4), since these do not depend on that estimate. It is perhaps illuminating to explain how the basic expression in (2) can be used in concrete situations, since this will illustrate the fact that N is not needed, and that the assumptions of WCR concerning this N are redundant. Expression (2) is valid in all circumstances where only a single individual S in a set D={d1,…,dn} of n profiles turns out not to be excluded as a candidate for being C. For the formal derivation of this result, it is irrelevant how large the database is, whether or not there was a suspicion against S, why the search was conducted, how the other individuals were excluded, or in which order evidence has been gathered. None of these additional aspects are needed to derive (2), and therefore they will not change (2) algebraically. But in order to assess the required probabilities, of course, these different situations can lead to different ‘numerical’ evaluations of these relevant probabilities, and therefore also to different numerical posterior odds. We next treat some special cases in more detail. For database searches, we distinguish between various different scenarios. In the first one, we assume that the search is carried out without any other relevant information about C other than the obtained profile, that is, there is no additional evidence against any database member. We call this a ‘cold case search’. After the search, we are informed that there is a match with a profile in the database, with or without knowing the identity of the donor of the matching profile. The second type of search, which we call a ‘targeted search’, arises if a suspect S has been identified and that suspect happens to already be in the database, say S = di. We then carry out the search to confirm this suspicion. In that case, a single match with another person would have surprised us much more than if we indeed obtain a single match with the already identified suspect. The classical ‘probable cause’ situation is the one where the identified suspect S is the only person whose profile is compared to that of C. Mathematically this corresponds to a targeted database search in a database consisting only of S. As we will see below, the size of the criminal population is never needed, and never relevant. In the former formulation in (6), we consider the hypotheses C = S versus C≠S⁠. In the latter formulation in (7), we consider the hypotheses C∈D versus C∉D⁠. In (6), we formulate hypotheses based on the database search ‘result’ and derive what the single match means for C = S; in (7), we formulate hypotheses based on the database ‘search’ itself and derive what the single match means for C∈D⁠. It remains to provide a numerical assessment of P(C∈D)⁠. One way to do so is to let P(C∈D) be equal to the proportion of traces that have been previously searched with and have given rise to a match in the database. In doing so, one implicitly assumes that all previous matches were with the true donor of the trace, and that the traces that were searched with in the past form a sufficiently representative sample to be useful for an estimate of P(C∈D)⁠. This assumption is not entirely unproblematic. If we take the type of crime into account, the estimate for P(C∈D) may change depending on whether the case is, for example, a burglary case, a homicide or a sexual assault case. Furthermore, one may argue that the trace donor C need not be the actual offender. This, however, may also be possible for previous searches; the probability P(C∈D) therefore applies to C as trace donor and not to C as offender. Bearing these cautions in mind, it is not uncommon for databases to be sufficiently large as to have odds P(C∈D)/P(C∉D) that are within one order of magnitude of being even. If that is the case, the posterior odds are of the same order of magnitude as the likelihood ratio, and we can then say that the odds on the match being with the trace donor are within one order of magnitude of 1/(np)⁠. If, for example, n = 106 and pS = 10−9, the odds are 1000: 1 that the match is with the actual trace donor. Of course, when the specifics of the crime and of the uncovered suspect are brought into consideration, these odds will need to be further updated. If, for example, it turns out that the match is with a person yet to be born when the crime was committed, they will be reduced to zero. But this cannot happen very often, since there will be a thousand true matches for every coincidental one for these n and p. In case, the match is indeed with the trace donor, and the trace donor is the actual offender, further evidence can potentially be uncovered, which will raise the odds from 1000: 1 to a larger number. When further non-genetic evidence I is found and taken into account, the result (2) still applies, but all probabilities need to be conditioned on I. This has no effect on the match probability pS but now P(C=S|I)>P(C=S)⁠. For D⁠, since additional evidence against one of its members S has been found, the probability that D contains C can not decrease, and hence we have P(C∉D|I)≤P(C∉D)⁠. Putting this together, we see that the posterior odds on C = S increase, reflecting the strengthening of the case against S due to the new evidence I. The preceding discussion brings us naturally to the targeted search case. In this case, evidence against S is found before the database search is done. Since there is no temporal order for probabilities, we must arrive at the same posterior odds regardless of whether S is identified via the database cold case search and further evidence is subsequently found, or when this happens in the reverse order. If we take into account the additional evidence before we process the evidence ES, we will no longer have P(C=S|C∈D)=1/n⁠, but a much larger value, approaching P(C=S|C∈D)≈1 as more and more evidence against S is uncovered. In that case, S was—before the database search—pretty much the only plausible candidate for C, which in turn means that P(C=S)≈P(C∈D)⁠, making the hypotheses C = S and C∈D much closer to being equivalent then in the cold case. In terms of (6) and (7), both terms P(C=S|C∈D) and P(C∉D|C≠S) are close to 1, so that the likelihood ratio is close to 1/p, regardless of whether we start out with hypotheses about S (in which case the likelihood ratio is larger than 1/p) or about D (in which case it is smaller). The exclusions that the database search has provided are, in other words, essentially irrelevant since we already believed that S was by far the most plausible candidate for being C before carrying out the search. Learning that the other database members, who we already believed not to be C, are indeed not C, then has only very little impact. Now we arrive naturally at the probable cause case, which we can think of in various ways. We can set D={S} so that no other comparisons have been done other than between S and C, who turned out to have matching profiles. Alternatively, we can think of a database in which all individuals apart from S were already excluded prior to the search, i.e. P(C=S|C∈D)=P(C∉D|C≠S)=1⁠. The latter formulation is nothing but an extreme case of the targeted search case, which we discussed above. Regardless of how we think about it, the hypotheses C = S and C∈D are then equivalent prior to learning ES, so that (6) and (7) coincide. The likelihood ratio in favour of C = S (or in favour of C∈D⁠, which is now the same hypothesis) is then exactly equal to 1/p. The reduction of the likelihood ratio of Wixted et al. (2019) to the expression (1), necessary for the analysis in Wixted et al. (2019), is too simplistic, as real cases are more complicated than that. The examples show that a setup in which population (sizes) play a role is not needed. All that matters is to evaluate the expression in (2). Thus, finally, we reiterate our position: the size of the active criminal population is not relevant for the calculation of the posterior odds. We do not dispute that w may under additional (strong!) assumptions be interpreted as related to that size as Wixted et al. (2019) do; our point is that this is a totally separate consideration, which is not relevant to the problem at hand. We consider the concept of active criminal population itself to be problematic and artificial, but the analysis in Wixted et al. (2019) can do without it. But—perhaps almost paradoxically—this in fact strengthens the conclusions of Wixted et al. (2019), since the paper effectively boils down to a proposal to measure the posterior odds, essentially using Stockmarr’s approach, from the so far obtained match rate in the database. Finally, we note that the current contribution has been concerned purely with the mathematical analysis of database searches. Our main intention is to steer the debate towards the relevant quantities only. We have not touched upon the question whether or not it is actually recommendable to report posterior odds, as we believe that this is a matter for debate outside the scope of this contribution.

Read the paper · More papers on PaperTik