Wednesday, 11 July 2012

More on Statistical Reasoning and the Higgs Boson

A posting of July 6, "The Probability that the Higgs Boson Has Been Discovered," mentioned the transposition of a p-value in stories in the popular press about the discovery of what is likely to be the Higgs Boson. Professor Dennis Lindley, a major figure in the development of Bayesian methods (and known to some readers of this blog as the author of a classic paper on using them to identify glass fragments) posed a few questions on the experiment via the list server of the International Society for Bayesian Analysis. One highly informed set of answers came from Louis Lyons (organiser of PHYSTAT series of meetings, and a member of CMS Collaboration at CERN). The following is a slightly edited version of the comments of Lindley (DL) and Lyons (LL). The comments presuppose knowledge of the meaning of a p-value, a likelihood ratio, Bayes' rule, and the divide between frequentists and Bayesians. (The original text as well as many other interesting messages are at http://bayesian.org/forums/news/3648.)

DL:
Specifically, the news referred to a confidence interval with 5-sigma limits.

LL:
The test statistic we use for looking at p-values is basically the likelihood ratio for the two hypotheses (H_0 = Standard Model (S. M.) of Particle Physics, but no Higgs; H_1 = S.M with Higgs). A small p_0 (and a reasonable p_1) then implies that H_1 is a better description of the data than H_0. This of course does not prove that H_1 is correct, but maybe Nature corresponds to some H_2, which is more like H_1 than it is like H_0. Indeed in principle data will never prove a theory is true, but the more experimental tests it survives, the happier we are to use it -- e.g. Newtonian mechanics was fine for centuries till the arrival of Relativity.

In the case of the Higgs, it can decay to different sets of particles, and these rates are defined by the S.M.  We measure these ratios, but with large uncertainties with the present data. They are consistent with the S.M. predictions, but it could be much more convincing with more data. Hence the caution about saying we have discovered the Higgs of the S.M.

DL:
Five standard deviations, assuming normality, means a p-value of around 0.0000005. A number of questions spring to mind.

1.  Why such an extreme evidence requirement? We know from a Bayesian perspective that this only makes sense if (a) the existence of the Higgs boson (or some other particle sharing some of its properties) has extremely small prior probability and/or (b) the consequences of erroneously announcing its discovery are dire in the extreme. Neither seems to be the case, so why 5-sigma?

LL:
This is an unfortunate tradition, that is used more readily by journal editors than by Particle Physicists. Reasons are
a) Historically we have had 3 and 4 sigma effects that have gone away

b) The 'Look Elsewhere Effect' (LEE). We are worried about the chance of a statistical fluctuation mimicking our observation, not only at the given mass of 125 GeV but anywhere in the spectrum. The quoted p-values are 'local' i.e. the chance of a fluctuation at the observed mass. Unfortunately the LEE correction factor is not very precisely defined, because of ambiguities about what is meant by 'elsewhere'

c) The possibility of some systematic effect (characterised by a nuisance parameter) being more important than allowed for in the analysis, or even overlooked - see the recent experiment at CERN which claimed that neutrinos travelled faster than the speed of light.

d) A subconscious use of Bayes Theorem to turn p-values into probabilities about the hypotheses.
All the above vary from experiment to experiment, so we realise that it is a bit unfair to use the same standard for discovery for all analyses. We prefer just to quote the p-values (or whatever).

DL:
2. Rather than ad hoc justification of a p-value, it is of course better to do a proper Bayesian analysis.  Are the particle physics community completely wedded to frequentist analysis?

LL:
No we are not anti-Bayesian, and indeed our test statistics is a likelihood ratio. If you like, you can regard our p-values as an attempt to calibrate the meaning of a particular value of the likelihood ratio.

We actually recommend that for parameter determination at the LHC, it is useful to compare Bayesian and Frequentist methods. But for comparing hypotheses (e.g. an experimental distribution is fitted by H_0 = a smooth distribution; or by H_1 = a smooth distribution plus a localised peak), we are worried about what priors to use for the extra parameters that occur in the alternative hypothesis.We would welcome advice.

DL:
3. We know that given enough data it is nearly always possible for a significance test to reject the null hypothesis at arbitrarily low p-values, simply because the parameter will never be exactly equal to its null value. And apparently the LHC has accumulated a very large quantity of data. So could even this extreme p-value be illusory?

LL:
We are aware of this. But in fact, although the LHC has accumulated enormous amounts of data, the Higgs search is like looking for a needle in  a haystack. The final samples of events that are used to look for the Higgs contain only tens to thousands of events.

These and related issues are discussed to some extent in my article "Open statistical issues in Particle Physics", Ann. Appl. Stat. Volume 2, Number 3 (2008), 887-915. It is supposed to be statistician-friendly.

Tuesday, 10 July 2012

If the Shoe Fits, You Must Not Calculate It (Part II)

The Court of Appeal in R. v. T. did not like Mr. Ryder’s truncated and standardized testimony (outlined in Part I). It complained about the supposedly small size of the FSS database, the recourse to a formula for combining information on different features, the mixture of objective and subjective probabilities in arriving at the undisclosed conditional probability of 1/100 and the unstated likelihood ratio of 100, the lack of the references in his testimony and reports to these computations and the FSS likelihood table, and the use of the honorific "scientific” in front of “evidence” and “support." In summarizing its reasons for judging the conviction "unsafe," the court emphasized the issue of transparency and completeness, writing that "the practice of using a Bayesian approach and likelihood ratios to formulate opinions placed before a jury without that process being disclosed and debated in court is contrary to principles of open justice."

Although Mr. Ryder insisted that he had merely used the figures to confirm what his "very extensive experience of footwear marks" already told him, the court saw the effort to reason more explicitly about the match as ammunition for cross-examination. Concluding that this cross-examination could have changed the outcome of the trial, the Court of Appeal quashed the conviction and ordered a retrial.

It suggested that on retrial, testimony not billed as scientific and based strictly on personal experience about the fact that the defendant’s Nike trainers “could have” been the source of the impressions would be acceptable. Beyond this, the court seemed willing to countenance "a more definite evaluative opinion" — as long as the "size or pattern" is "unusual" based on "years studying this kind of comparison." That kind of opinion would be fine, the court wrote, because "[i]t is a judgment based on his experience"  and "without any figures or mathematical formula."

There is much to criticize in this court’s reasoning. As many authors have noted, surely an expert whose intuitive or experiential impressions give rise to a judgment about the source hypothesis should be encouraged to consult the available statistical data and to consider their limitations to produce a fully informed judgment.

Less prominent in the writing on the case is the fact that, despite a description to the court, from Visiting Professor and Scientist and Scholar Allan Jamieson, of Mr. Ryder's "approach as 'the Bayesian approach' of using likelihood ratios," likelihoodism and Bayesianism are hardly the same. As explained in the previous posting (Part I), Mr. Ryder never spoke of prior or posterior odds or of source probabilities. He merely attached some words—"modest support"—to his unstated estimate of the likelihood ratio.The court's condemnation of "the practice of using a Bayesian approach" therefore seems inapposite.

To be sure, "Bayesianism" can be used to motivate the likelihood ratio as a measure of probative value, but that does not make the simple presentation of a likelihood ratio “the Bayesian approach.” In fact, the likelihood school of statistical inference abjures the use of prior probability distributions, and it does not use Bayes' rule in coming to decisions about hypotheses. Likelihoodism maintains that the statistician should be concerned only with whether the evidence provides increased or decreased support for one hypothesis over another. A likelihoodist would find the presentation of the likelihood ratio itself, without any Bayesian baggage or interpretation, entirely appropriate.

Of course, this is not to say that either the likelihoodist or the Bayesian would agree that the particular likelihood ratio kept out of sight in R. v. T. should be admissible. Although the likelihoodist would be pleased that Mr. Ryder's LR of 100 and his description of it as "moderate ... support" were untainted by a subjective, prior probability, the objectivity or accuracy of the estimate and the adjective could be a source of legitimate concern.

But the concern is not with the use of likelihoods per se. Casting doubt on a particular estimate of an LR does not make it appropriate for the expert to speak of the source probability—quantitatively or qualitatively. Indeed, if the expert lacks the data and experience with which to estimate the likelihood ratio, as the court in R. v. T. suggested, how can the expert have anything useful to say about the source probability? The court's preference for expert opinions on source probabilities simply sweeps the problem under the proverbial rug.

References on R. v. T.

C.E.H. Berger et al., Evidence Evaluation: A Response to the Court of Appeal Judgment in R v T, 51 Sci. & Justice 43 (2011)

F. Hoar et al., Extending the Confusion about Bayes, 74 Modern L. Rev. 444 (2011)

David H. Kaye, Likelihoodism, Bayesianism, and a Pair of Shoes, 53 Jurimetrics J. (forthcoming Fall 2012)

G.S. Morrison, The Likelihood-ratio Framework and Forensic Evidence in Court: A Response to R v T, 16 Int'l J. Evid. & Proof 1 (2012)

Mike Redmayne et al., Forensic Science Evidence in Questions, 2011 Crim. L.R. 347

References on Likelihoodism

Jeffrey D. Blume, Likelihood Methods for Measuring Statistical Evidence, 21 Stat. Med. 2563 (2002)

Anthony W.F. Edwards, Likelihood (2d ed. 1992)

James Hawthorne, Inductive Logic, in Stanford Encyclopedia of Philosophy (Edward N. Zalta ed. 2012)

Richard M. Royall, Statistical Evidence: a Likelihood Paradigm (1997)

Monday, 9 July 2012

If the Shoe Fits, You Must Not Calculate It (Part I)

In R. v. T., [2010] EWCA Crim. 2439, the Court of Appeal of England and Wales wrote an opinion that dismayed, if not enraged, leading forensic scientists across the globe. The brouhaha began with testimony in a murder trial that there was "a moderate degree of scientific evidence to support the view that the [Nike trainers recovered from the appellant] had made the footwear marks."

This evidence came from "Mr Ryder of the Forensic Science Service (FSS)." Mr. Ryder compared four aspects of the footwear marks from a murder scene and a pair of Nike "trainers found in the appellant's house after his arrest," namely:
  • Pattern (p). The FSS maintained a database of the characteristics of the shoes it inspected. About 20% had the pattern of the soles of the Nikes—the same pattern seen in the shoeprints. The probability of the pattern in a pair of shoes not worn at the crime scene (-W) would be P(p | -W) = 1/5, where the vertical line stands for "given" or "conditional on."
  • Size (s). According to another database, 3% of shoes sold with that pattern were size 11 (UK). Given the uncertainty in the precise size of a shoe that might have left the marks and in the effects of wear, the examiner adjusted this last figure upward. He estimated that as many as 10% of shoes sold would be in the right size range. Hence, P(p & s | -W) = 1/10).
  • Wear (w). He estimated (somehow) that about 50% of relevant shoes would show as much wear as was indicated by the impression and the shoes themselves. P(p & s & w | -W) = 1/2.
  • Damage (d). Finally, he felt that he that marks indicative of damage to the shoes added almost nothing to the other information. P(p & s & w & d | -W) = 1.
It follows that if the marks did not come from the defendant's shoes, the probability that they would be comparable to the ones at the murder scene in these four respects would be P(p & s & w & d | -W) = (1/5)(1/10)(1/2)(1) = 1/100.

A frequentist statistician might say that the similarities between the impressions and the defendant’s shoes are good evidence that the shoes left the marks because a p-value of 0.01 is small.

A likelihoodist statistician would want to know more. It is not enough to believe that an outcome is improbable under the defense’s hypothesis that the defendant’s shoes did not leave the marks (-W). One also must consider the probability of the marks under the prosecution’s hypothesis that the defendant’s shoes left the marks (W). The "law of likelihood" postulates that when the probability of the evidence under one hypothesis exceeds that under the competing, simple hypothesis, it supports the former over the latter to a degree given by the ratio of the conditional probabilities. If the two probabilities in the "likelihood ratio" are equal, then the evidence is to be expected to the same extent under both hypotheses. It cannot help us discriminate between them. Thus, some law review article writers have called the likelihood ratio a "relevance ratio."

Here, the probability that the impressions would match the shoes if they had indeed come from the defendant's Nike trainers was almost 100%, so Mr. Ryder concluded that the evidence (E = p & s & w & d) was about 100 times more probable if the marks came from the defendant's shoes (W) than if they came from other shoes (-W). In symbols, the likelihood ratio (LR) for his conditional probabilities is


LR = P(E | W) / P(E | -W) = 1 / (1/100) = 100.

Mr. Ryder made this rough estimate "to confirm an opinion substantially based on his experience and so that it could be expressed in a standardised form." He wrote three reports and testified, but never once did he mention these numbers. Rather, he testified that "In my opinion there is a moderate degree of scientific support the view that the [Nike trainers] made those marks."

He chose the word "moderate" from a table that the Forensic Science Service had selected for ranges of the likelihood ratio. The table, which he did not mention at trial or in his written reports, classified LRs from 10-100 as providing “moderate support.” The use of a standard table of "verbal equivalents" finds approval in reports of the European Association of Forensic Service Providers and a committee of the US National Research Council.

A Bayesian statistician would agree that a likelihood ratio of 100 supports the prosecution's theory substantially more than the defense's. But this statistician would not stop here. He would argue that the LR is a "Bayes factor." It raises the prior odds on W by 100. A juror willing to post prior odds of only 1 to 10 for the prosecution's hypothesis before hearing Mr. Ryder's evidence (and harboring no doubts about the veracity and accuracy of that evidence) now should be willing to revise the odds upward. Specifically, Bayes' rule gives posterior odds of LR x prior odds = 100 x 1/10 = 10 to 1. Whatever the value V of the prior odds, the posterior odds for this evidence are 100V.

Mr. Ryder stopped with the verbiage derived from the likelihoods and the FSS table. He did not give a Bayesian interpretation to the evidence -- something that the Court of Appeal had strongly disapproved of in earlier cases. Even so, the court in R. v. T. held that his testimony of "a moderate degree of scientific evidence to support the [prosecution's] view" rendered the conviction unsafe and therefore required a new trial.The next posting on the topic will explain why.

Saturday, 7 July 2012

Have DNA Databases Produced False Convictions?

About a year ago, I asked whether any false convictions have resulted from DNA database searches [1]. Of course, if there are any, they might be hard to find, but there is a known recent case of a false initial accusation. It came about because a laboratory contaminated a crime-scene sample with DNA from an whose DNA was on file from other cases.

In March 2012, a private firm in England re-used a "plastic tray[] as part of the robotic DNA extraction process" [2]. The tray, which should have been disposed of, apparently contained some DNA from Adam Scott, a young man from Exeter, in Devon [3]. This DNA contaminated the sample from the clothing of a woman who had been raped in a park in Manchester. Police charged Scott, who vehemently protested that he had never been to Manchester, with the rape. After detectives realized that Scott "was in prison 300 miles away, awaiting trial on other unrelated offences" at the time of the rape, the charges were dropped [3]. An audit and investigation of 26,000 other samples analyzed after the robotic system had been introduced, uncovered no other instances of contamination. Steps intended to prevent a repetition of the error have been implemented [3].

Other errors in handling samples have been documented. In a 2001 Las Vegas case, police obtained DNA samples from two young suspects, Dwayne Jackson and his cousin, Howard Grissom. A technician put Jackson's sample in a vial marked as Grissom's, and vice versa. A falsely accused Jackson then pleaded guilty and was imprisoned for four years. The error came to light in 2010, after Grissom was convicted of robbing and stabbing a woman in Southern California. California officials took Grissom's DNA and entered the profile into the national database, leading to a match to the crime-scene DNA from the 2001 burglary for which Jackson had been falsely convicted [4].

Of course, this is not a case of a DNA database hit producing a conviction or even a false accusation. Quite the contrary, it is a case of a DNA database producing an exoneration that would not have occurred otherwise. But both cases vividly illustrate the need to implement quality control systems that reduce the chance of handling and other errors and to avoid over-reliance on cold hits.

Added July 9, 2012

Jeremy Gans's comment, posted this morning, is required reading.

References

1. David H. Kaye, Genetic Justice: Potential and Real, The Double Helix Law Blog, June 5, 2011.

2. BBC News, DNA Blunder: Man Accused of Rape After Human Error, Mar. 21, 2012.

3. Simon Israel, DNA Contamination Blamed on Human Error, Channel 4 News, May 9, 2012.

4. Lawrence Mower & Doug McMurdo, Las Vegas Police Reveal DNA Error Put Wrong Man in Prison, Las Vegas Rev.-J., July 8, 2011.

Cross-posted from The Double Helix Law Blog.

Friday, 6 July 2012

The Probability that the Higgs Boson Has Been Discovered

Surely everyone has heard of the probable discovery of the Higgs boson. But what does it have to do with  forensic science or law? It is a reminder that the "prosecutor's fallacy" is not limited to prosecutors or courtrooms. Reports in the popular press by skilled physicists and science writers trying to explain this impressive discovery are replete with a messy form of the transposition fallacy. Here is an example from an otherwise excellent report by physicist Lawrence Krauss in Slate magazine:
One can in fact quantify the likelihood that the observations are mistaken and that the events are actually background noise mimicking a real signal. Each experiment quotes a likelihood of very close to “5 sigma,” meaning the likelihood that the events were produced by chance is less than one in 3.5 million. Yet in spite of this, the only claim that has been made so far is that the new particle is real and “Higgs-like.”
Likewise, Nature announced "just a 0.00006% probability that the result is due to chance." The New York Times reported that "the likelihood that their signal was a result of a chance fluctuation was less than one chance in 3.5 million, 'five sigma,' which is the gold standard in physics for a discovery," attributing the statement to CERN's physicists.

How is this (mis)reporting related to the transposition fallacy? Well, sigma (σ) stands for standard deviation, and 5σ means 5 standard deviations from the value expected if the measurements were just noise. For a normal distribution, results this extreme or more extreme would be seen in pure noise a small fraction of the time. The tiny figures quoted above are estimates of that fraction. The fraction is the statistician's p-value, P(>5σ | noise), and it is on the order of 10-6. In plain English (and one bit of Greek), the probability of data of more than 5σ given that they are just noise is on the order of one in a million. So the observations would be very surprising if they were just noise.

But the probability that they actually are noise is an inverse probability, P(noise | data). That probability depends on the likelihoods P(5σ | noise) and P(5σ | signal) as well as on the prior probability, P(noise). The p-value itself does not generally "quantify the likelihood that the observations are mistaken and that the events are actually background noise mimicking a real signal." It does not specify the "probability that the result is due to chance." If one wants to quantify the probability that the data are a real signal rather than noise, then, for better or worse, one must turn to Bayes' rule.

References (for physicists)

- Giulio D’Agostini, Bayesian Reasoning in High Energy Physics, CERN Yellow Report 99-03, July 1999
- Giulio D'Agostini, Probability and Measurement Uncertainty in Physics: A Bayesian Primer (1995)

A couple of other blogs (and one newspaper) making the same point

- http://www.r-bloggers.com/the-higgs-boson-sigma-5-and-the-concept-of-p-values/
- http://understandinguncertainty.org/higgs-it-one-sided-or-two-sided
- http://randomastronomy.wordpress.com/2012/07/04/higgs-boson-discovery-and-how-to-not-interpret-p-values/
- http://blog.carlislerainey.com/2012/07/07/innumeracy-and-higgs-boson/
- http://understandinguncertainty.org/explaining-5-sigma-higgs-how-well-did-they-do#comment-1449

- http://online.wsj.com/article/SB10001424052702303962304577509213491189098.html

Postscript

Professor Dennis Lindley, a major figure in the development of Bayesian methods (and known to some readers of this blog as the author of a classic paper on using them to identify glass fragments) posed a few questions on the Higgs boson experiment via the list server of the International Society for Bayesian Analysis. One well informed set of answers came from Louis Lyons (organiser of PHYSTAT series of meetings, and a member of CMS Collaboration at CERN). I posted a slightly edited version on July 11 under the title "More on Statistical Reasoning and the Higgs Boson." The full text of these and various other interesting messages is at http://bayesian.org/forums/news/3648.

Thursday, 28 June 2012

The Arizona Supreme Court Adopts a No-Peeking Rule for Juvenile Arrestee DNA

Preface: This posting replaces one from June 28. Part of that initial discussion of the Arizona Supreme Court's opinion was, I think, unwarranted. In particular the criticism of the court's treatment of the state interests may not have been accurate. Complex opinions, like good literature, rarely can be fully grasped on a first reading.
* * *
A few days ago, the Supreme Court of Arizona promulgated a creative “don’t peek” rule for DNA samples routinely taken from juveniles before a finding of delinquency. Justice Andrew Hurwitz (who has just moved to the U.S. Court of Appeals for the Ninth Circuit) penned the unanimous opinion in Mario W. v. Kaipio, Commissioner, No. CV-11-0344-PR (Ariz. June 27, 2012). The opinion injects some new ideas and analysis into the legal controversy over arrestee DNA sampling, but I have to question whether the reasoning is sufficient to support the result the court reaches and to ask how far the court's theory of Fourth Amendment privacy extends.

At the outset, the Arizona court quite properly sets out the normal rule that Fourth Amendment reasonableness requires a warrant and probable cause unless a categorical exception to these requirements exists. But then the court states that “[t]he parties do not dispute the applicability of the totality of the circumstances test, and we therefore analyze the Arizona scheme under that rubric.” This is hardly a ringing endorsement of this mode of analysis, but it is the way most courts approach the issue [1].

Getting to the specifics of DNA sampling on arrest, the court observes that there are “two separate intrusions” and “two searches — ‘the physical collection of the DNA sample’ and the ‘processing of the DNA sample.’” The former observation is basically correct. “The seizure of buccal cells is a physical intrusion, but does not reveal by itself intimate personal information about the individual.”1/

But the laboratory analysis probably is not a “later search.” The U.S. Supreme Court, at any rate, has yet to hold that physical testing or inspection is a separate search simply because it produces information about the substance being analyzed. Indeed, the Court, in two opinions—United States v. Edwards, 415 U.S. 800 (1974), and United States v. Jacobsen, 466 U.S. 109 (1984)—has held the opposite.

The Arizona court relies on an analogy to containers. It maintains that human cells are like steamer trunks or purses that contain private possessions. The police engage in a search when they open such a container and rummage through its contents.

The analogy looks good at first blush. People surely have reasonable expectations of privacy in the contents of their luggage and their purses. The Orthodox Jew on Yom Kippur with an apple core in her purse, the Catholic juvenile with birth control pills in hers, and the English literature professor with sleazy novels in his trunk all have a fair claim to freedom from unregulated intrusions into their purses or luggage. The police will all but inevitably espy these legal but embarrassing items if they look through the container without a warrant.

But compare this with the laboratory analysis of the epithelial cells. The laboratory extracts a single kind of molecule—DNA. It does not look at the rest of the cell. Within the DNA, it looks at a tiny fraction of the genome—locations (“loci”) that are not potentially embarrassing (except insofar as they match crime-scene samples).1/ The situation begins to resemble cases in which dogs that (supposedly) alert only to drugs are used to sniff luggage—and that, the Court has twice held, is not a search.2/

Because the government does not look through the parts of genome in which an individual has a strong expectation of privacy, a better analogy is required. Imagine, then, that every time a person commits a crime, a mysterious being delivers an envelope to the police that always contains only two things—a card with the name of an individual who was at the scene of the crime (but not necessarily at the time the crime occurred) and a key to a safe deposit box in that person’s name. Is opening the envelope a “search” that triggers the need for a warrant or an exception to the warrant requirement? Maybe, but the cases and the doctrine cited in Mario W. are insufficient to establish this result. All that the container cases establish is that the police must abide by the constitutional requirements for searches before and when they use the key to open the safe deposit box. The box, of course, is the vast part of the human genome that the police do not open in DNA testing for identity. In DNA profiling for law enforcement databases, they only read the name on the card.3/

Yet, whether one denominates the laboratory analysis as a separate search is not decisive. It might be a constitutionally permissible, warrantless, probable-causeless search, at least under the totality-of-the-circumstances balancing test. The Arizona justices reject this conclusion in favor of the following rule: (1) the state’s “important interest in locating an absconding juvenile and, perhaps years after charges were filed, ascertaining that the person located is the one previously charged” justifies collecting the sample—“even if a formal judicial determination of probable cause was not made at the advisory hearing.” However, (2) no combination of state interests justifies the warrantless laboratory analysis of the DNA sample (a) to determine whether it matches unsolved crime samples or (b) to have a profile in a database that will identify the juvenile as the contributor of DNA found in future crimes.

But why is taking DNA solely for “locating an absconding juvenile” so critical when the state already takes fingerprints that can be used this purpose? Doesn't the fingerprint on file eliminate the need to house the DNA as well, as the Maryland Court of Appeals recently held in King v. State, 42 A.3d 549 (Md. 2012)?

The Arizona court’s answer is that “[o]ne arrested for a serious crime may be fingerprinted before a judicial determination of probable cause. ... A judicial order to provide a buccal cell sample occasions no constitutionally distinguishable intrusion.” This suggests that the state can choose either fingerprints or DNA as the source of identifying marks. However, if a DNA profile is “intimate personal information about the individual” merely because it constitutes “uniquely identifying information”—which is all that Mario W. says about informational privacy—then fingerprints are equally “intimate personal information.” They too provide “uniquely identifying information.” Indeed, they are better for this purpose, for they permit differentiation of identical twins.

So does Mario W. prohibit the state from examining the minutiae in fingerprints unless or until arrestees are convicted (the no-peeking rule)? From running an arrestee’s print against a database of prints from unsolved crimes? From adding the fingerprint to the national Automated Fingerprint Identification System database (AFIS) before that point? Of course, DNA loci might be significantly more threatening to privacy than fingerprint details, but that conclusion is far from obvious [1].

In analyzing the state’s interests in pre-conviction DNA analysis, the opinion correctly notes that the value in solving unrelated crimes (and in deterring future ones) is reduced considerably by two features of the Arizona law. As with all pre-conviction profiling and databasing, many of the arrestees would have their samples analyzed and included in the state database after they are convicted anyway. As for the ones who are not convicted, the Arizona law does not permit continued use of the profiles. Thus, the opinion notes, with current technology and staffing, the government has the benefit of the profiles for only a month or so (for those who not adjudicated delinquent) and for only an extra month or so for the others.

These points help explain the court's balancing, but how enduring are they? Advances in technology, making it possible to analyze profiles in a matter of hours, easily could extend the period of pre-conviction use. In addition, what would happen if the law did not require the samples to be removed from the database in the event that the state does not prove delinquency? Obviously, that would advance the state's law enforcement interests (although it might not be politically popular). The sad fact is that lots of people who are arrested but never convicted commit later crimes. If DNA is to be believed, California's "Grim Sleeper" killer is one. Lives would have been saved had his profile been acquired at his first of sixteen arrests and kept in a state database. Of course, the mere fact that law enforcement could gain by keeping tabs on more people cannot make all such practices constitutional. Still, post-adjudication retention of juvenile DNA profiles in all cases adds something to the state's interests that is missing in the Arizona system of juvenile arrestee DNA databasing and that therefore must be considered in totality balancing for that more extensive system.

Interestingly, the Mario W. court intimates that expungement is mandatory "given the constitutional presumption of innocence" and the fact that those accused of crimes "do not forfeit Fourth Amendment protections." This part of the opinion raises several puzzles. Given the history and cases on the presumption of innocence, it is an expansive reading of the presumption [2]. Moreover, if the presumption does mean that the state may not include DNA profiles of those arrested but not ultimately convicted in databases, what of fingerprints, which are retained indefinitely? As noted earlier, the court's theory as to why DNA profiling invades informational privacy seems to apply with equal force to AFIS databases. That people to not forfeit Fourth Amendment rights just because they are accused of crimes—or, for that matter, convicted of them—is important, but it does not imply that the Fourth Amendment is an absolute barrier to suspicionless profiling and databasing. The opinion asserts that

[O]ne accused of a crime, although having diminished expectations of privacy in some respects, does not forfeit Fourth Amendment protections with respect to other offenses not charged absent either probable cause or reasonable suspicion. An arrest for vehicular homicide, for example, cannot alone justify a warrantless search of an arrestee’s financial records to see if he is also an embezzler.

As with the purse and the trunk, the financial records of the arrestee merit strong Fourth Amendment protection (unless, according the U.S. Supreme Court, they are held by a bank or other third party). But what is it about the DNA loci that merits similar protection? The state’s claim is not that an arrest justifies every unrelated search. It is (or should be) that the custodial arrest justifies using identifying marks—whether they are within fingerprint impressions or DNA molecules—for identification of the person and then for speculative searching against the marks left at past and future unsolved crimes.

To be sure, a sensitive balancing of individual and public interests might lead to the conclusion that the latter goes too far. But the assumption in Mario W. seems to be that tokens of an individual’s identity are necessarily “intimate personal information” that impose a "serious intrusion on ... privacy interests." Without a clearer and more convincing analysis of the actual privacy interests associated with the many things that mark us as individuals—DNA profiles, fingerprints, iris scans, even photographs—Mario W. raises more questions than it answers.

Notes

1. One can quibble with the term “seizure,” for the extraction of the cells in the inner surface of the cheek does not seem to be a seizure in the Fourth Amendment sense. Unlike keeping a person away from his home or luggage or stopping him, it is not a substantial interference with the individual’s use of his possessions or his person. It is, however, probably a search under Cupp v. Murphy, 412 U.S. 291 (1973) (physical intrusion under fingernail), Schmerber v. California, 384 U.S. 757 (1966) (physical intrusion with syringe), or United States v. Jones, No. 10–1259 (U.S. Jan. 23, 2012) (the GPS tracking case that applied a trespass-with-intent-to-acquire-information test for ascertaining a “search”).

2. But Caballes and Place also are distinguishable in that DNA loci are not contraband.)

3. How much, if any, other information the card contains is an interesting question.

4. I am oversimplifying. When profiles from putative close relatives are available, the loci can be used for kinship testing. For example, if the state has the profiles of a mother-father-child trio, it could determine whether they are in the specified biological relationship or whether, for example, someone else is the biological father. The reader is invited to make his own comparison between the strength of the privacy interest in the contents of all manner of containers of personal effects and records on the one hand, and the STR loci used for identification, on the other.

References

1. David H. Kaye, A Fourth Amendment Theory for Arrestee DNA and Other Biometric Databases, 15 U. Pa. J. Const. L. No. 4 (forthcoming Apr. 2013).

2. David H. Kaye, Drawing Lines: Unrelated Probable Cause as a Prerequisite to Early DNA Collection, 91 N. C. L. Rev. Addendum No. 1 (forthcoming Oct. 2012)

Postscript

Rereading the Mario W. opinion yet again, the following paragraph struck me:
¶26 The State argues that once it has lawfully obtained the cell samples, the Fourth Amendment provides no greater bar to the processing of those samples and the extraction of the DNA profile than it does to the analysis of fingerprints. But the State's reliance on the fingerprinting analogy here is misplaced. Once fingerprints are obtained, no further intrusion on the privacy of the individual is required before they can be used for investigative purposes. In this sense, the fingerprint is akin to a photograph or voice exemplar. But before DNA samples can be used by law enforcement, they must be physically processed and a DNA profile extracted. See Erin Murphy, The New Forensics: Criminal Justice, False Certainty, and the Second Generation of Scientific Evidence, 95 Cal. L. Rev. 721, 726-30 (2007).
This is a distinction without a difference. First, in both fingerprinting and DNA analysis, a sample (an exemplar) must be collected from an arrestee. Elsewhere, the opinion describes the intrusion on the individual in this step with unusual clarity. Second, with both fingerprinting and DNA profiling, the physical sample must be examined "before [the] samples can be used by law enforcement."

The fingerprint information lies in minutiae that must be studied by eye or by computer to extract useful data. The DNA information lies in particular loci that must be characterized by chemical reactions and computers to extract useful data. What matters is not the physics or the chemistry, but the transformation into identifying information. If extracting this information is a separate search for DNA, then extracting the identifying information also is a separate search for fingerprints. If this "second search" requires a warrant for DNA, it requires it for fingerprints.

Cross-posted from Double Helix Law

Sunday, 24 June 2012

Fingerprinting Error Rates Down Under

How accurate is latent fingerprint identification? Considering that fingerprints are the most common form of trace evidence (for crimes in general), this is a vital question. The other day, I mentioned one court’s view that a false identification rate of 0.01% is only marginally higher than 1 in 11 million. The former statistic comes from a well designed experiment—the Noblis-FBI study published in 2011.\1/ The latter comes from a U.S. Court of Appeals opinion citing the earlier testimony of an FBI fingerprint supervisor.

A few months after the publication of the Noblis-FBI study, a short report of a second controlled experiment on the ability of latent print examiners to match those prints to exemplars appeared, this time in the journal Psychological Science.\2/ University of Queensland psychology lecturer Jason M. Tangen and two co-authors recruited 37 “qualified practicing fingerprint experts from five police organizations” and 37 college students to see how they would do in evaluating pairs of prints—some from the same fingers (mates) and some from different fingers (nonmates). Id. at 995.

The Task

Although the Australian researchers wrote that the task they gave their subjects “emulates the most forensically relevant aspect of the identification process,” id., the professional examiners and the students did not have the options of declaring a pair of images unsuitable for analysis or ultimately inconclusive.\3/ Instead, these “[p]articipants were asked to judge whether the prints in each pair matched, using a confidence rating scale ranging from 1 (sure different) to 12 (sure same) ... [with] ratings of 1 through 6 indicat[ing] a match [and] ratings of 7 through 12 indicat[ing] no match.” Id.

The members of the two groups each received “36 simulated crime-scene prints that were paired with fully rolled prints.” Id. Some of the nonmates were supposed to be similar to the latent print of the pair because they were the closest matches (according to a computer program) in the Australian National Automated Fingerprint Identification System (ANAFIS). The other nonmates were plucked at random from the research database and ANAFIS. Each participant received a 12 latent prints paired with “similar” nonmates, 12 paired with “nonsimilar” nonmates, and 12 paired with mates. Id. at 996.

How They Did

False negative rate. The “experts performed exceedingly well.” Id. at 997. In the 12 × 37 = 444 trials of mates, “experts correctly identified 92.12% of the pairs, on average, as matches (hits), misidentifying 7.88% as nonmatches (misses).” Id.

False positive rate. For the pairs intended to be difficult (the similar nonmates), “experts correctly declared nearly all of the pairs (99.32%) to be nonmatches (correct rejections); only 3 pairs (0.68%) out of the 444 in this condition were incorrectly declared to be matches (false alarms).” Id. at 997. Furthermore, not a single expert “misidentified any of the 12 nonsimilar distractor prints as matches.” Id.

Students. The undergraduates did not fare nearly as well as the practitioners. For example, they “mistakenly identified 55.18% of the [pairs of] similar ... [nonmates] as matches.” Id.

The following two tables list more or less comparable error rates from both studies.\4/


Table 1. False negative rates

Professionals
(US)
Professionals
(Australia)
Undergraduates
Mates 450/4113
(10.9%)
35/444
(7.88%)
113/444
(25.45%)


Table 2. False positive rates

Professionals
(US)
Professionals
(Australia)
Undergraduates
Similar nonmates 6/3628
(0.02%)
3/444
(0.68%)
245/444
(55.18%)
Nonsimilar nonmates 0/444
(0.00%)
102/444
(22.97%)

Discussion

As both sets of researchers appreciate, these rates do not necessarily generalize to casework. They simply exemplify what can be achieved under the experimental conditions, in which the subjects knew they were being tested. Nevertheless, the sensitivity and specificity of professional examiners in ascertaining when a pair of prints emanates from the same finger should help mute the most extreme criticism of the field. By the same token, they should prompt investigations of the conditions under which misclassifications tend to occur and, as Tangen et al. note, they “should affect the testimony of forensic examiners and the assertions that they can reasonably make.” Id. at 997.

The Australian researchers are impressed by the extent to which professional examiners outperformed undergraduate students. But some of the gap could be caused in part by a difference in motivation. The students received course credit for turning in answers, but they may have had little incentive to agonize over the best classification in every one of the 36 comparisons they were asked to make. This confounding variable should be considered before making unequivocal claims of “a real performance benefit” that “may satisfy legal admissibility criteria.” Id.

Notes

1. Bradford T. Ulery, R. Austin Hicklin, JoAnn Buscaglia & Maria Antonia Roberts, Accuracy and Reliability of Forensic Latent Fingerprint Decisions, 108 Proc. Nat’l Acad. Sci. 7733 (2011), available at http://www.pnas.org/content/108/19/7733.full.pdf.

2. Jason M. Tangen, Matthew B. Thompson, and Duncan J. McCarthy, Identifying Fingerprint Expertise, 22 Psych. Sci. 995 (2011), available at http://mbthompson.com/wp-content/uploads/2011/03/TangenThompsonMcCarthyIdentifyingFingerprintExpertisePsycScience2011.pdf.

3. The third author, a latent print analyst, believed that all the simulated latent prints were of value for identification. Id. at 996.

4. The studies differ in significant ways, limiting the number of comparisons that one can make and the confidence that one can have in direct comparisons of the statistics. The Noblis-FBI study used no randomly selected nonmate exemplars—all the nonmate pairs were “similar” within the meaning of the Australian study. Also, only practicing fingerprint examiners participated, so the student-professional dichotomy has not been replicated. To account as much as possible for the fact that all the pairs of prints had to compared and an ultimate conclusion had to be drawn in the Australian experiment, the rates from the U.S. study are based on instances in which the subjects deemed the prints to be of value for individualization and reached a conclusion about the comparison. Even if this adjustment is adequate for some purposes, however, it does not eliminate the possibility that the inability to use the “inconclusive” category contributed to the larger false positive rate of the Australian subjects.