Monday, 9 July 2012

If the Shoe Fits, You Must Not Calculate It (Part I)

In R. v. T., [2010] EWCA Crim. 2439, the Court of Appeal of England and Wales wrote an opinion that dismayed, if not enraged, leading forensic scientists across the globe. The brouhaha began with testimony in a murder trial that there was "a moderate degree of scientific evidence to support the view that the [Nike trainers recovered from the appellant] had made the footwear marks."

This evidence came from "Mr Ryder of the Forensic Science Service (FSS)." Mr. Ryder compared four aspects of the footwear marks from a murder scene and a pair of Nike "trainers found in the appellant's house after his arrest," namely:
  • Pattern (p). The FSS maintained a database of the characteristics of the shoes it inspected. About 20% had the pattern of the soles of the Nikes—the same pattern seen in the shoeprints. The probability of the pattern in a pair of shoes not worn at the crime scene (-W) would be P(p | -W) = 1/5, where the vertical line stands for "given" or "conditional on."
  • Size (s). According to another database, 3% of shoes sold with that pattern were size 11 (UK). Given the uncertainty in the precise size of a shoe that might have left the marks and in the effects of wear, the examiner adjusted this last figure upward. He estimated that as many as 10% of shoes sold would be in the right size range. Hence, P(p & s | -W) = 1/10).
  • Wear (w). He estimated (somehow) that about 50% of relevant shoes would show as much wear as was indicated by the impression and the shoes themselves. P(p & s & w | -W) = 1/2.
  • Damage (d). Finally, he felt that he that marks indicative of damage to the shoes added almost nothing to the other information. P(p & s & w & d | -W) = 1.
It follows that if the marks did not come from the defendant's shoes, the probability that they would be comparable to the ones at the murder scene in these four respects would be P(p & s & w & d | -W) = (1/5)(1/10)(1/2)(1) = 1/100.

A frequentist statistician might say that the similarities between the impressions and the defendant’s shoes are good evidence that the shoes left the marks because a p-value of 0.01 is small.

A likelihoodist statistician would want to know more. It is not enough to believe that an outcome is improbable under the defense’s hypothesis that the defendant’s shoes did not leave the marks (-W). One also must consider the probability of the marks under the prosecution’s hypothesis that the defendant’s shoes left the marks (W). The "law of likelihood" postulates that when the probability of the evidence under one hypothesis exceeds that under the competing, simple hypothesis, it supports the former over the latter to a degree given by the ratio of the conditional probabilities. If the two probabilities in the "likelihood ratio" are equal, then the evidence is to be expected to the same extent under both hypotheses. It cannot help us discriminate between them. Thus, some law review article writers have called the likelihood ratio a "relevance ratio."

Here, the probability that the impressions would match the shoes if they had indeed come from the defendant's Nike trainers was almost 100%, so Mr. Ryder concluded that the evidence (E = p & s & w & d) was about 100 times more probable if the marks came from the defendant's shoes (W) than if they came from other shoes (-W). In symbols, the likelihood ratio (LR) for his conditional probabilities is


LR = P(E | W) / P(E | -W) = 1 / (1/100) = 100.

Mr. Ryder made this rough estimate "to confirm an opinion substantially based on his experience and so that it could be expressed in a standardised form." He wrote three reports and testified, but never once did he mention these numbers. Rather, he testified that "In my opinion there is a moderate degree of scientific support the view that the [Nike trainers] made those marks."

He chose the word "moderate" from a table that the Forensic Science Service had selected for ranges of the likelihood ratio. The table, which he did not mention at trial or in his written reports, classified LRs from 10-100 as providing “moderate support.” The use of a standard table of "verbal equivalents" finds approval in reports of the European Association of Forensic Service Providers and a committee of the US National Research Council.

A Bayesian statistician would agree that a likelihood ratio of 100 supports the prosecution's theory substantially more than the defense's. But this statistician would not stop here. He would argue that the LR is a "Bayes factor." It raises the prior odds on W by 100. A juror willing to post prior odds of only 1 to 10 for the prosecution's hypothesis before hearing Mr. Ryder's evidence (and harboring no doubts about the veracity and accuracy of that evidence) now should be willing to revise the odds upward. Specifically, Bayes' rule gives posterior odds of LR x prior odds = 100 x 1/10 = 10 to 1. Whatever the value V of the prior odds, the posterior odds for this evidence are 100V.

Mr. Ryder stopped with the verbiage derived from the likelihoods and the FSS table. He did not give a Bayesian interpretation to the evidence -- something that the Court of Appeal had strongly disapproved of in earlier cases. Even so, the court in R. v. T. held that his testimony of "a moderate degree of scientific evidence to support the [prosecution's] view" rendered the conviction unsafe and therefore required a new trial.The next posting on the topic will explain why.

Saturday, 7 July 2012

Have DNA Databases Produced False Convictions?

About a year ago, I asked whether any false convictions have resulted from DNA database searches [1]. Of course, if there are any, they might be hard to find, but there is a known recent case of a false initial accusation. It came about because a laboratory contaminated a crime-scene sample with DNA from an whose DNA was on file from other cases.

In March 2012, a private firm in England re-used a "plastic tray[] as part of the robotic DNA extraction process" [2]. The tray, which should have been disposed of, apparently contained some DNA from Adam Scott, a young man from Exeter, in Devon [3]. This DNA contaminated the sample from the clothing of a woman who had been raped in a park in Manchester. Police charged Scott, who vehemently protested that he had never been to Manchester, with the rape. After detectives realized that Scott "was in prison 300 miles away, awaiting trial on other unrelated offences" at the time of the rape, the charges were dropped [3]. An audit and investigation of 26,000 other samples analyzed after the robotic system had been introduced, uncovered no other instances of contamination. Steps intended to prevent a repetition of the error have been implemented [3].

Other errors in handling samples have been documented. In a 2001 Las Vegas case, police obtained DNA samples from two young suspects, Dwayne Jackson and his cousin, Howard Grissom. A technician put Jackson's sample in a vial marked as Grissom's, and vice versa. A falsely accused Jackson then pleaded guilty and was imprisoned for four years. The error came to light in 2010, after Grissom was convicted of robbing and stabbing a woman in Southern California. California officials took Grissom's DNA and entered the profile into the national database, leading to a match to the crime-scene DNA from the 2001 burglary for which Jackson had been falsely convicted [4].

Of course, this is not a case of a DNA database hit producing a conviction or even a false accusation. Quite the contrary, it is a case of a DNA database producing an exoneration that would not have occurred otherwise. But both cases vividly illustrate the need to implement quality control systems that reduce the chance of handling and other errors and to avoid over-reliance on cold hits.

Added July 9, 2012

Jeremy Gans's comment, posted this morning, is required reading.

References

1. David H. Kaye, Genetic Justice: Potential and Real, The Double Helix Law Blog, June 5, 2011.

2. BBC News, DNA Blunder: Man Accused of Rape After Human Error, Mar. 21, 2012.

3. Simon Israel, DNA Contamination Blamed on Human Error, Channel 4 News, May 9, 2012.

4. Lawrence Mower & Doug McMurdo, Las Vegas Police Reveal DNA Error Put Wrong Man in Prison, Las Vegas Rev.-J., July 8, 2011.

Cross-posted from The Double Helix Law Blog.

Friday, 6 July 2012

The Probability that the Higgs Boson Has Been Discovered

Surely everyone has heard of the probable discovery of the Higgs boson. But what does it have to do with  forensic science or law? It is a reminder that the "prosecutor's fallacy" is not limited to prosecutors or courtrooms. Reports in the popular press by skilled physicists and science writers trying to explain this impressive discovery are replete with a messy form of the transposition fallacy. Here is an example from an otherwise excellent report by physicist Lawrence Krauss in Slate magazine:
One can in fact quantify the likelihood that the observations are mistaken and that the events are actually background noise mimicking a real signal. Each experiment quotes a likelihood of very close to “5 sigma,” meaning the likelihood that the events were produced by chance is less than one in 3.5 million. Yet in spite of this, the only claim that has been made so far is that the new particle is real and “Higgs-like.”
Likewise, Nature announced "just a 0.00006% probability that the result is due to chance." The New York Times reported that "the likelihood that their signal was a result of a chance fluctuation was less than one chance in 3.5 million, 'five sigma,' which is the gold standard in physics for a discovery," attributing the statement to CERN's physicists.

How is this (mis)reporting related to the transposition fallacy? Well, sigma (σ) stands for standard deviation, and 5σ means 5 standard deviations from the value expected if the measurements were just noise. For a normal distribution, results this extreme or more extreme would be seen in pure noise a small fraction of the time. The tiny figures quoted above are estimates of that fraction. The fraction is the statistician's p-value, P(>5σ | noise), and it is on the order of 10-6. In plain English (and one bit of Greek), the probability of data of more than 5σ given that they are just noise is on the order of one in a million. So the observations would be very surprising if they were just noise.

But the probability that they actually are noise is an inverse probability, P(noise | data). That probability depends on the likelihoods P(5σ | noise) and P(5σ | signal) as well as on the prior probability, P(noise). The p-value itself does not generally "quantify the likelihood that the observations are mistaken and that the events are actually background noise mimicking a real signal." It does not specify the "probability that the result is due to chance." If one wants to quantify the probability that the data are a real signal rather than noise, then, for better or worse, one must turn to Bayes' rule.

References (for physicists)

- Giulio D’Agostini, Bayesian Reasoning in High Energy Physics, CERN Yellow Report 99-03, July 1999
- Giulio D'Agostini, Probability and Measurement Uncertainty in Physics: A Bayesian Primer (1995)

A couple of other blogs (and one newspaper) making the same point

- http://www.r-bloggers.com/the-higgs-boson-sigma-5-and-the-concept-of-p-values/
- http://understandinguncertainty.org/higgs-it-one-sided-or-two-sided
- http://randomastronomy.wordpress.com/2012/07/04/higgs-boson-discovery-and-how-to-not-interpret-p-values/
- http://blog.carlislerainey.com/2012/07/07/innumeracy-and-higgs-boson/
- http://understandinguncertainty.org/explaining-5-sigma-higgs-how-well-did-they-do#comment-1449

- http://online.wsj.com/article/SB10001424052702303962304577509213491189098.html

Postscript

Professor Dennis Lindley, a major figure in the development of Bayesian methods (and known to some readers of this blog as the author of a classic paper on using them to identify glass fragments) posed a few questions on the Higgs boson experiment via the list server of the International Society for Bayesian Analysis. One well informed set of answers came from Louis Lyons (organiser of PHYSTAT series of meetings, and a member of CMS Collaboration at CERN). I posted a slightly edited version on July 11 under the title "More on Statistical Reasoning and the Higgs Boson." The full text of these and various other interesting messages is at http://bayesian.org/forums/news/3648.

Thursday, 28 June 2012

The Arizona Supreme Court Adopts a No-Peeking Rule for Juvenile Arrestee DNA

Preface: This posting replaces one from June 28. Part of that initial discussion of the Arizona Supreme Court's opinion was, I think, unwarranted. In particular the criticism of the court's treatment of the state interests may not have been accurate. Complex opinions, like good literature, rarely can be fully grasped on a first reading.
* * *
A few days ago, the Supreme Court of Arizona promulgated a creative “don’t peek” rule for DNA samples routinely taken from juveniles before a finding of delinquency. Justice Andrew Hurwitz (who has just moved to the U.S. Court of Appeals for the Ninth Circuit) penned the unanimous opinion in Mario W. v. Kaipio, Commissioner, No. CV-11-0344-PR (Ariz. June 27, 2012). The opinion injects some new ideas and analysis into the legal controversy over arrestee DNA sampling, but I have to question whether the reasoning is sufficient to support the result the court reaches and to ask how far the court's theory of Fourth Amendment privacy extends.

At the outset, the Arizona court quite properly sets out the normal rule that Fourth Amendment reasonableness requires a warrant and probable cause unless a categorical exception to these requirements exists. But then the court states that “[t]he parties do not dispute the applicability of the totality of the circumstances test, and we therefore analyze the Arizona scheme under that rubric.” This is hardly a ringing endorsement of this mode of analysis, but it is the way most courts approach the issue [1].

Getting to the specifics of DNA sampling on arrest, the court observes that there are “two separate intrusions” and “two searches — ‘the physical collection of the DNA sample’ and the ‘processing of the DNA sample.’” The former observation is basically correct. “The seizure of buccal cells is a physical intrusion, but does not reveal by itself intimate personal information about the individual.”1/

But the laboratory analysis probably is not a “later search.” The U.S. Supreme Court, at any rate, has yet to hold that physical testing or inspection is a separate search simply because it produces information about the substance being analyzed. Indeed, the Court, in two opinions—United States v. Edwards, 415 U.S. 800 (1974), and United States v. Jacobsen, 466 U.S. 109 (1984)—has held the opposite.

The Arizona court relies on an analogy to containers. It maintains that human cells are like steamer trunks or purses that contain private possessions. The police engage in a search when they open such a container and rummage through its contents.

The analogy looks good at first blush. People surely have reasonable expectations of privacy in the contents of their luggage and their purses. The Orthodox Jew on Yom Kippur with an apple core in her purse, the Catholic juvenile with birth control pills in hers, and the English literature professor with sleazy novels in his trunk all have a fair claim to freedom from unregulated intrusions into their purses or luggage. The police will all but inevitably espy these legal but embarrassing items if they look through the container without a warrant.

But compare this with the laboratory analysis of the epithelial cells. The laboratory extracts a single kind of molecule—DNA. It does not look at the rest of the cell. Within the DNA, it looks at a tiny fraction of the genome—locations (“loci”) that are not potentially embarrassing (except insofar as they match crime-scene samples).1/ The situation begins to resemble cases in which dogs that (supposedly) alert only to drugs are used to sniff luggage—and that, the Court has twice held, is not a search.2/

Because the government does not look through the parts of genome in which an individual has a strong expectation of privacy, a better analogy is required. Imagine, then, that every time a person commits a crime, a mysterious being delivers an envelope to the police that always contains only two things—a card with the name of an individual who was at the scene of the crime (but not necessarily at the time the crime occurred) and a key to a safe deposit box in that person’s name. Is opening the envelope a “search” that triggers the need for a warrant or an exception to the warrant requirement? Maybe, but the cases and the doctrine cited in Mario W. are insufficient to establish this result. All that the container cases establish is that the police must abide by the constitutional requirements for searches before and when they use the key to open the safe deposit box. The box, of course, is the vast part of the human genome that the police do not open in DNA testing for identity. In DNA profiling for law enforcement databases, they only read the name on the card.3/

Yet, whether one denominates the laboratory analysis as a separate search is not decisive. It might be a constitutionally permissible, warrantless, probable-causeless search, at least under the totality-of-the-circumstances balancing test. The Arizona justices reject this conclusion in favor of the following rule: (1) the state’s “important interest in locating an absconding juvenile and, perhaps years after charges were filed, ascertaining that the person located is the one previously charged” justifies collecting the sample—“even if a formal judicial determination of probable cause was not made at the advisory hearing.” However, (2) no combination of state interests justifies the warrantless laboratory analysis of the DNA sample (a) to determine whether it matches unsolved crime samples or (b) to have a profile in a database that will identify the juvenile as the contributor of DNA found in future crimes.

But why is taking DNA solely for “locating an absconding juvenile” so critical when the state already takes fingerprints that can be used this purpose? Doesn't the fingerprint on file eliminate the need to house the DNA as well, as the Maryland Court of Appeals recently held in King v. State, 42 A.3d 549 (Md. 2012)?

The Arizona court’s answer is that “[o]ne arrested for a serious crime may be fingerprinted before a judicial determination of probable cause. ... A judicial order to provide a buccal cell sample occasions no constitutionally distinguishable intrusion.” This suggests that the state can choose either fingerprints or DNA as the source of identifying marks. However, if a DNA profile is “intimate personal information about the individual” merely because it constitutes “uniquely identifying information”—which is all that Mario W. says about informational privacy—then fingerprints are equally “intimate personal information.” They too provide “uniquely identifying information.” Indeed, they are better for this purpose, for they permit differentiation of identical twins.

So does Mario W. prohibit the state from examining the minutiae in fingerprints unless or until arrestees are convicted (the no-peeking rule)? From running an arrestee’s print against a database of prints from unsolved crimes? From adding the fingerprint to the national Automated Fingerprint Identification System database (AFIS) before that point? Of course, DNA loci might be significantly more threatening to privacy than fingerprint details, but that conclusion is far from obvious [1].

In analyzing the state’s interests in pre-conviction DNA analysis, the opinion correctly notes that the value in solving unrelated crimes (and in deterring future ones) is reduced considerably by two features of the Arizona law. As with all pre-conviction profiling and databasing, many of the arrestees would have their samples analyzed and included in the state database after they are convicted anyway. As for the ones who are not convicted, the Arizona law does not permit continued use of the profiles. Thus, the opinion notes, with current technology and staffing, the government has the benefit of the profiles for only a month or so (for those who not adjudicated delinquent) and for only an extra month or so for the others.

These points help explain the court's balancing, but how enduring are they? Advances in technology, making it possible to analyze profiles in a matter of hours, easily could extend the period of pre-conviction use. In addition, what would happen if the law did not require the samples to be removed from the database in the event that the state does not prove delinquency? Obviously, that would advance the state's law enforcement interests (although it might not be politically popular). The sad fact is that lots of people who are arrested but never convicted commit later crimes. If DNA is to be believed, California's "Grim Sleeper" killer is one. Lives would have been saved had his profile been acquired at his first of sixteen arrests and kept in a state database. Of course, the mere fact that law enforcement could gain by keeping tabs on more people cannot make all such practices constitutional. Still, post-adjudication retention of juvenile DNA profiles in all cases adds something to the state's interests that is missing in the Arizona system of juvenile arrestee DNA databasing and that therefore must be considered in totality balancing for that more extensive system.

Interestingly, the Mario W. court intimates that expungement is mandatory "given the constitutional presumption of innocence" and the fact that those accused of crimes "do not forfeit Fourth Amendment protections." This part of the opinion raises several puzzles. Given the history and cases on the presumption of innocence, it is an expansive reading of the presumption [2]. Moreover, if the presumption does mean that the state may not include DNA profiles of those arrested but not ultimately convicted in databases, what of fingerprints, which are retained indefinitely? As noted earlier, the court's theory as to why DNA profiling invades informational privacy seems to apply with equal force to AFIS databases. That people to not forfeit Fourth Amendment rights just because they are accused of crimes—or, for that matter, convicted of them—is important, but it does not imply that the Fourth Amendment is an absolute barrier to suspicionless profiling and databasing. The opinion asserts that

[O]ne accused of a crime, although having diminished expectations of privacy in some respects, does not forfeit Fourth Amendment protections with respect to other offenses not charged absent either probable cause or reasonable suspicion. An arrest for vehicular homicide, for example, cannot alone justify a warrantless search of an arrestee’s financial records to see if he is also an embezzler.

As with the purse and the trunk, the financial records of the arrestee merit strong Fourth Amendment protection (unless, according the U.S. Supreme Court, they are held by a bank or other third party). But what is it about the DNA loci that merits similar protection? The state’s claim is not that an arrest justifies every unrelated search. It is (or should be) that the custodial arrest justifies using identifying marks—whether they are within fingerprint impressions or DNA molecules—for identification of the person and then for speculative searching against the marks left at past and future unsolved crimes.

To be sure, a sensitive balancing of individual and public interests might lead to the conclusion that the latter goes too far. But the assumption in Mario W. seems to be that tokens of an individual’s identity are necessarily “intimate personal information” that impose a "serious intrusion on ... privacy interests." Without a clearer and more convincing analysis of the actual privacy interests associated with the many things that mark us as individuals—DNA profiles, fingerprints, iris scans, even photographs—Mario W. raises more questions than it answers.

Notes

1. One can quibble with the term “seizure,” for the extraction of the cells in the inner surface of the cheek does not seem to be a seizure in the Fourth Amendment sense. Unlike keeping a person away from his home or luggage or stopping him, it is not a substantial interference with the individual’s use of his possessions or his person. It is, however, probably a search under Cupp v. Murphy, 412 U.S. 291 (1973) (physical intrusion under fingernail), Schmerber v. California, 384 U.S. 757 (1966) (physical intrusion with syringe), or United States v. Jones, No. 10–1259 (U.S. Jan. 23, 2012) (the GPS tracking case that applied a trespass-with-intent-to-acquire-information test for ascertaining a “search”).

2. But Caballes and Place also are distinguishable in that DNA loci are not contraband.)

3. How much, if any, other information the card contains is an interesting question.

4. I am oversimplifying. When profiles from putative close relatives are available, the loci can be used for kinship testing. For example, if the state has the profiles of a mother-father-child trio, it could determine whether they are in the specified biological relationship or whether, for example, someone else is the biological father. The reader is invited to make his own comparison between the strength of the privacy interest in the contents of all manner of containers of personal effects and records on the one hand, and the STR loci used for identification, on the other.

References

1. David H. Kaye, A Fourth Amendment Theory for Arrestee DNA and Other Biometric Databases, 15 U. Pa. J. Const. L. No. 4 (forthcoming Apr. 2013).

2. David H. Kaye, Drawing Lines: Unrelated Probable Cause as a Prerequisite to Early DNA Collection, 91 N. C. L. Rev. Addendum No. 1 (forthcoming Oct. 2012)

Postscript

Rereading the Mario W. opinion yet again, the following paragraph struck me:
¶26 The State argues that once it has lawfully obtained the cell samples, the Fourth Amendment provides no greater bar to the processing of those samples and the extraction of the DNA profile than it does to the analysis of fingerprints. But the State's reliance on the fingerprinting analogy here is misplaced. Once fingerprints are obtained, no further intrusion on the privacy of the individual is required before they can be used for investigative purposes. In this sense, the fingerprint is akin to a photograph or voice exemplar. But before DNA samples can be used by law enforcement, they must be physically processed and a DNA profile extracted. See Erin Murphy, The New Forensics: Criminal Justice, False Certainty, and the Second Generation of Scientific Evidence, 95 Cal. L. Rev. 721, 726-30 (2007).
This is a distinction without a difference. First, in both fingerprinting and DNA analysis, a sample (an exemplar) must be collected from an arrestee. Elsewhere, the opinion describes the intrusion on the individual in this step with unusual clarity. Second, with both fingerprinting and DNA profiling, the physical sample must be examined "before [the] samples can be used by law enforcement."

The fingerprint information lies in minutiae that must be studied by eye or by computer to extract useful data. The DNA information lies in particular loci that must be characterized by chemical reactions and computers to extract useful data. What matters is not the physics or the chemistry, but the transformation into identifying information. If extracting this information is a separate search for DNA, then extracting the identifying information also is a separate search for fingerprints. If this "second search" requires a warrant for DNA, it requires it for fingerprints.

Cross-posted from Double Helix Law

Sunday, 24 June 2012

Fingerprinting Error Rates Down Under

How accurate is latent fingerprint identification? Considering that fingerprints are the most common form of trace evidence (for crimes in general), this is a vital question. The other day, I mentioned one court’s view that a false identification rate of 0.01% is only marginally higher than 1 in 11 million. The former statistic comes from a well designed experiment—the Noblis-FBI study published in 2011.\1/ The latter comes from a U.S. Court of Appeals opinion citing the earlier testimony of an FBI fingerprint supervisor.

A few months after the publication of the Noblis-FBI study, a short report of a second controlled experiment on the ability of latent print examiners to match those prints to exemplars appeared, this time in the journal Psychological Science.\2/ University of Queensland psychology lecturer Jason M. Tangen and two co-authors recruited 37 “qualified practicing fingerprint experts from five police organizations” and 37 college students to see how they would do in evaluating pairs of prints—some from the same fingers (mates) and some from different fingers (nonmates). Id. at 995.

The Task

Although the Australian researchers wrote that the task they gave their subjects “emulates the most forensically relevant aspect of the identification process,” id., the professional examiners and the students did not have the options of declaring a pair of images unsuitable for analysis or ultimately inconclusive.\3/ Instead, these “[p]articipants were asked to judge whether the prints in each pair matched, using a confidence rating scale ranging from 1 (sure different) to 12 (sure same) ... [with] ratings of 1 through 6 indicat[ing] a match [and] ratings of 7 through 12 indicat[ing] no match.” Id.

The members of the two groups each received “36 simulated crime-scene prints that were paired with fully rolled prints.” Id. Some of the nonmates were supposed to be similar to the latent print of the pair because they were the closest matches (according to a computer program) in the Australian National Automated Fingerprint Identification System (ANAFIS). The other nonmates were plucked at random from the research database and ANAFIS. Each participant received a 12 latent prints paired with “similar” nonmates, 12 paired with “nonsimilar” nonmates, and 12 paired with mates. Id. at 996.

How They Did

False negative rate. The “experts performed exceedingly well.” Id. at 997. In the 12 × 37 = 444 trials of mates, “experts correctly identified 92.12% of the pairs, on average, as matches (hits), misidentifying 7.88% as nonmatches (misses).” Id.

False positive rate. For the pairs intended to be difficult (the similar nonmates), “experts correctly declared nearly all of the pairs (99.32%) to be nonmatches (correct rejections); only 3 pairs (0.68%) out of the 444 in this condition were incorrectly declared to be matches (false alarms).” Id. at 997. Furthermore, not a single expert “misidentified any of the 12 nonsimilar distractor prints as matches.” Id.

Students. The undergraduates did not fare nearly as well as the practitioners. For example, they “mistakenly identified 55.18% of the [pairs of] similar ... [nonmates] as matches.” Id.

The following two tables list more or less comparable error rates from both studies.\4/


Table 1. False negative rates

Professionals
(US)
Professionals
(Australia)
Undergraduates
Mates 450/4113
(10.9%)
35/444
(7.88%)
113/444
(25.45%)


Table 2. False positive rates

Professionals
(US)
Professionals
(Australia)
Undergraduates
Similar nonmates 6/3628
(0.02%)
3/444
(0.68%)
245/444
(55.18%)
Nonsimilar nonmates 0/444
(0.00%)
102/444
(22.97%)

Discussion

As both sets of researchers appreciate, these rates do not necessarily generalize to casework. They simply exemplify what can be achieved under the experimental conditions, in which the subjects knew they were being tested. Nevertheless, the sensitivity and specificity of professional examiners in ascertaining when a pair of prints emanates from the same finger should help mute the most extreme criticism of the field. By the same token, they should prompt investigations of the conditions under which misclassifications tend to occur and, as Tangen et al. note, they “should affect the testimony of forensic examiners and the assertions that they can reasonably make.” Id. at 997.

The Australian researchers are impressed by the extent to which professional examiners outperformed undergraduate students. But some of the gap could be caused in part by a difference in motivation. The students received course credit for turning in answers, but they may have had little incentive to agonize over the best classification in every one of the 36 comparisons they were asked to make. This confounding variable should be considered before making unequivocal claims of “a real performance benefit” that “may satisfy legal admissibility criteria.” Id.

Notes

1. Bradford T. Ulery, R. Austin Hicklin, JoAnn Buscaglia & Maria Antonia Roberts, Accuracy and Reliability of Forensic Latent Fingerprint Decisions, 108 Proc. Nat’l Acad. Sci. 7733 (2011), available at http://www.pnas.org/content/108/19/7733.full.pdf.

2. Jason M. Tangen, Matthew B. Thompson, and Duncan J. McCarthy, Identifying Fingerprint Expertise, 22 Psych. Sci. 995 (2011), available at http://mbthompson.com/wp-content/uploads/2011/03/TangenThompsonMcCarthyIdentifyingFingerprintExpertisePsycScience2011.pdf.

3. The third author, a latent print analyst, believed that all the simulated latent prints were of value for identification. Id. at 996.

4. The studies differ in significant ways, limiting the number of comparisons that one can make and the confidence that one can have in direct comparisons of the statistics. The Noblis-FBI study used no randomly selected nonmate exemplars—all the nonmate pairs were “similar” within the meaning of the Australian study. Also, only practicing fingerprint examiners participated, so the student-professional dichotomy has not been replicated. To account as much as possible for the fact that all the pairs of prints had to compared and an ultimate conclusion had to be drawn in the Australian experiment, the rates from the U.S. study are based on instances in which the subjects deemed the prints to be of value for individualization and reached a conclusion about the comparison. Even if this adjustment is adequate for some purposes, however, it does not eliminate the possibility that the inability to use the “inconclusive” category contributed to the larger false positive rate of the Australian subjects.

Friday, 22 June 2012

Fingerprint Error Rates in Love

In United States v. Love, No. 10cr2418–MMM, 2011 WL 2173644 (S.D. Cal. June 1, 2011), the “government charged Donny Love, Sr. with crimes related to the May 4, 2008 bombing of the federal courthouse in San Diego. ... Before trial, Love moved to exclude the testimony of Robin Ruth, the standards and practices program manager” of the FBI's latent fingerprint unit. Id. at *1.

Love argued that latent fingerprint analysis lacks the scientific or other foundation for admissibility under the federal rules of evidence. The district judge, M. Margaret McKeown, held a pretrial hearing at which Ruth, who "possesses degrees in genetic engineering and in forensic and biological anthropology," id. at *9, was the only witness. In her (inexplicably unreported) opinion, Judge McKeown described the testimony on the risk of error as follows:
Latent fingerprint examiners have sometimes stated that their analyses have a zero error rate. . . . Ruth did not so testify. Rather, she stated that the ACE–V methodology is unbiased and that the methodology itself introduces no random error. Errors occur, but those errors are human errors resulting from human implementation of the ACE–V process. . . . Because human errors are nonsystematic, Ruth believes that there is no overall predictive error rate in latent fingerprint analysis.
Id. at *5. This sounds like gobbledy-gook to me. To begin with, the FBI’s continued insistence that there is an ACE-V “methodology” separate from the human being who is doing the analysis, comparison, and evaluation is mystifying. As a NIST expert working group on human factors in fingerprinting recently wrote, human errors are a function of the system that includes human beings. Such a system can be structured to favor errors of one type—false identifications—or the other—false eliminations. The Noblis-FBI study (see postings of April 26, 27, 30, and May 1) indicates that in the analysis phase, examiners systematically discard prints that contain useful information. In the evaluation phase, they systematically err in favor of false eliminations over false identifications.

Moreover, the reasoning that “[b]ecause human errors are nonsystematic, ... there is no overall predictive error rate in latent fingerprint analysis” is hard to decipher. Long-term features of random (nonsystematic) processes are predictable. If a laboratory’s examiners make false identifications and false eliminations, each with a constant probability of 0.001, then the “human errors are nonsystematic” but the 0.001 error rate can be used to predict the incidence of errors in casework. Ms. Ruth probably was saying that error probabilities are not constant and that existing data do not permit very realistic or accurate statistical modeling of errors in a laboratory.

In any event, a better argument is available. It is clear that fingerprint examinations can produce useful information—they can make correct source attributions at levels far greater than chance alone would produce. The opinion refers to
very few . . . cases in which an examiner identifies a latent print as matching a known print even though the two prints were actually made by different individuals. Most significantly, the May 2011 [Noblis-FBI] study of the performance of 169 fingerprint examiners revealed a total of six false positives among 4,083 comparisons of non-matching fingerprints for “an overall false positive rate of 0.1%.” See Ulery et al., supra, at 7733, 7735.
Id. Of course, the 0.1% figure from a controlled experiment in which the examiners knew they were being tested might not be “predictive” of the rates in casework. An experiment reveals what happens under the conditions in the experiment. Generalizing to other situations takes judgment. Apparently, Ms. Ruth thought that the error rate in practice would be smaller “because the prints used in the study were of relatively low quality.” Id. And, the court noted that verification by a second examiner should reduce the rate still more. Id.

The observed error rate of 0.1%, the court wrote, was only “marginally higher than the rates the FBI has previously estimated.” Id. The only previous statistic that the court cites is “1 in 11 million cases.” Id. The difference between 1/1,000 and 1/11,000,000 strikes me as more serious than the term “marginal” connotes. The government ought to devote more resources to reducing the risk of error if the chance of a false identification is as high as 1/1,000 than if it is a mere 1/11,000,000. The former may be "very low," but the latter is incredibly low.

Furthermore, historically derived error statistics such as 1/11,000,000 are all highly problematic in the absence of any mechanism that would uncover possible errors in run-of-the-mill cases. Consequently, and contrary to the view expressed in the opinion, one should not regard a supervisor's recollections of errors and the outcomes of investigations of occasional complaints of misidentifications as "evidence suggest[ing] that the rate is much lower than that figure [of 0.1%]." Id. at *6.

Appendix

Love is the first case to generate a widely available opinion discussing the Noblis-FBI study. It is reproduced below:

United States v. Love, No. 10cr2418–MMM, 2011 WL 2173644 (S.D. Cal. June 1, 2011) (not reported)

ORDER REGARDING DONNY LOVE, SR.'S MOTION TO EXCLUDE LATENT FINGERPRINT TESTIMONY

Honorable M. MARGARET McKEOWN, District Judge.

The government charged Donny Love, Sr. with crimes related to the May 4, 2008 bombing of the federal courthouse in San Diego. A trial on those charges began on May 23, 2011. Before trial, Love moved to exclude the testimony of Robin Ruth, the standards and practices program manager of Federal Bureau of Investigation's (“FBI”) latent fingerprint unit. Love argued that latent fingerprint analysis generally, and therefore Ruth's specific testimony about fifteen latent prints she analyzed for this case, are insufficiently reliable for admission under Federal Rule of Evidence 702 and the Supreme Court's opinions in Daubert v. Merrell Dow Pharmaceuticals, 509 U.S. 579 (1993), and Kumho Tire Co. v. Carmichael, 526 U.S. 137 (1999). At Love's request, on May 16, 2011, the court held a Daubert hearing regarding the latent fingerprint evidence. Ruth testified at that hearing regarding latent fingerprint analysis generally, the FBI's practices, and the analyses she performed in the course of her work on this case. No other witnesses were called to testify at the hearing. Because the May 16 hearing took place one week before trial, in order to provide sufficient notice of the ruling, the court provided an oral ruling at the close of the hearing and denied Love's motion to exclude the latent fingerprint testimony. The court also gave a brief explanation of that ruling, but stated that it would issue a written order elaborating on its reasons for denying Love's motion. This order provides that further explanation. To the extent anything in this order differs from the court's oral ruling of May 16, 2011, this order supersedes the oral ruling.

I. Latent Fingerprint Analysis

There are, broadly speaking, two kinds of fingerprints. “Rolled,” “full,” or “known” prints are taken under controlled circumstances, often for official use. “Latent” prints, by contrast, are left—often unintentionally—on the surfaces of objects touched by a person. Latent fingerprint examiners compare unidentified latent prints to known prints as a means of identifying the person who left the latent print.

In United States v. Mitchell, 365 F.3d 215 (3d Cir.2004), the Third Circuit provided a succinct introduction to the terminology used by latent fingerprint examiners and to the methodology employed by FBI examiners. In that court's words,

[f]ingerprints are left by the depositing of oil upon contact between a surface and the friction ridges of fingers. The field uses the broader term “friction ridge” to designate skin surfaces with ridges evolutionarily adapted to produce increased friction (as compared to smooth skin) for gripping. Thus toeprint or handprint analysis is much the same as fingerprint analysis. The structure of friction ridges is described in the record before us at three levels of increasing detail, designated as Level 1, Level 2 and Level 3. Level 1 detail is visible with the naked eye; it is the familiar pattern of loops, arches, and whorls. Level 2 detail involves “ridge characteristics”--the patterns of islands, dots, and forks formed by the ridges as they begin and end and join and divide. The points where ridges terminate or bifurcate are often referred to as “Galton points,” whose eponym, Sir Francis Galton, first developed a taxonomy for these points. The typical human fingerprint has somewhere between 75 and 175 such ridge characteristics. Level 3 detail focuses on microscopic variations in the ridges themselves, such as the slight meanders of the ridges (the “ridge path”) and the locations of sweat pores. This is the level of detail most likely to be obscured by distortions.
 

The FBI ... uses an identification method known as ACE–V, an acronym for “analysis, comparison, evaluation, and verification.” The basic steps taken by an examiner under this protocol are first to winnow the field of candidate matching prints by using Level 1 detail to classify the latent print. Next, the examiner will analyze the latent print to identify Level 2 detail (i.e., Galton points and their spatial relationship to one another), along with any Level 3 detail that can be gleaned from the print. The examiner then compares this to the Level 2 and Level 3 detail of a candidate full-rolled print (sometimes taken from a database of fingerprints, sometimes taken from a suspect in custody), and evaluates whether there is sufficient similarity to declare a match.

Id. at 221–22. At the evaluation step, the examiner may reach three conclusions—that the prints are a match (“identification”), that the prints are not a match (“exclusion”), or that more information is needed (“inconclusive”). See Hearing Ex. 1, at 42. “In the final step, the match is independently verified by another examiner.” Mitchell 365 F.3d at 222. It bears noting that since Mitchell was decided 2004, there have been important refinements and developments with respect to latent print identification, as documented in testimony before the court.

II. General Admissibility of Latent Fingerprint Evidence

As with all scientific or technical evidence, latent fingerprint evidence is admissible in court if it “will assist the trier of fact to understand or to determine a fact in issue” and “if (1) the testimony is based upon sufficient facts or data, (2) the testimony is the product of reliable principles and methods, and (3) the witness has applied the principles and methods reliably to the facts of the case.” Fed.R.Evid. 702. “Many factors ... bear on the [reliability] inquiry.” Daubert, 509 U.S. at 592–93. The Supreme Court has identified several factors that may often be relevant. These include whether the “technique can be (and has been) tested,” “[w]hether it has been subjected to peer review and publication,” the “known or potential rate of error,” “whether there are standards controlling the technique's operation,” and “whether the ... technique enjoys general acceptance within a relevant scientific community.” Kumho Tire, 526 U.S. at 149–50 (internal quotation marks and alteration omitted); accord Daubert, 509 U.S. at 593–94. Nevertheless, “the test of reliability is flexible, and Daubert 's list of specific factors neither necessarily nor exclusively applies to all experts or in every case.” Kumho, 526 U.S. at 141 (internal quotation marks omitted).

“The test under Daubert is not the correctness of the expert's conclusions but the soundness of [her] methodology.” Primiano v. Cook, 598F.3d558, 564 (9th Cir.2010). As a result, “the role of the courts in reviewing proposed expert testimony is to analyze expert testimony in the context of its field to determine if it is acceptable science.” Boyd v. City & Cnty. of S.F., 576 F.3d 938. 946 (9th Cir.2009). So long as “an expert meets the threshold established by Rule 702 as explained in Daubert, the expert may testify and the jury decides how much weight to give that testimony.” Primiano, 598 F.3d at 565. Put another way, although the court “must ensure that ‘junk science’ plays no part in the [jury's] decision,” Elsayed Mukhtar v. Cal. State Univ., Hayward, 299 F.3d 1053, 1063 (9th Cir.2002), “[s]haky but admissible evidence is to be attacked by cross examination, contrary evidence, and attention to the burden of proof, not exclusion,” Primiano, 598 F.3d at 564.

In this case, the parties discuss each of the five factors mentioned in Daubert and Kumho, They also address two additional factors drawn from United States v. Downing, 753 F.2d 1224 (3d Cir.1985)—namely, whether latent fingerprint analysis has a “relationship to ... established modes of scientific analysis” and whether it has “non-judicial uses.” Id. at 1238–39. The court addresses each factor in turn.

A. Testing

The first factor discussed in Daubert and Kumho asks whether a methodology can be tested and whether it has been tested. The parties do not dispute that the reliability of latent fingerprint analysis can be tested, and the record reveals three categories of potential tests. First, latent fingerprint analysis rests on two hypotheses—that fingerprints are unique to an individual and that prints persist, which is to say that they do not change over the course of an individual's life.\1/ These hypotheses are open to testing. The uniqueness hypothesis would be falsified if two people were found to share identical fingerprints, while the persistence hypothesis would be falsified if the same finger produced prints with details that evolved over time. Second, it is possible, at least in principle, to learn “the prevalence of different ridge flows and crease patterns” and “ridge formations and clusters of ridge formations” across individuals. See National Research Council of the National Academies, Strengthening Forensic Science in the United States: A Path Forward 144 (2009) (“NAS Report”), available at: http:// www.nap.edu/catalog/12589.html. This information would facilitate estimates of the reliability of the conclusion that two specific fingerprints are from the same individual. See Opp'n, Ex. B, at 8. Third, it is possible to test the reliability of conclusions reached by a given fingerprint analyst through controlled examinations.

1. At the hearing, Ruth testified that permanent scars alter an individual's fingerprints and thus are a known exception to the persistence hypothesis.

The fact that latent fingerprint analysis can be tested for reliability, without more, allows the first Daubert “factor to weigh in support of admissibility.” See Mitchell, 365 F.3d at 238. Ruth also testified, however, that at least some actual testing and research has been performed along each of the three dimensions discussed above. Testing of the uniqueness and persistence hypotheses dates to the eighteenth century and includes, among other things, a simple longitudinal study performed by Sir William Galton in the 1890s, a 1982 study of twins, and contemporary Bayesian statistical models. See Hearing Ex. 1, at 26–28; see also Mitchell, 365 F.3d at 236 & n. 16 (discussing a test of 50,000 fingerprints for uniqueness). Some recent statistical models also bear on the distribution of particular ridge characteristics across the population as a whole. See Opp'n, Ex. B, at 8 (citing three additional studies). Finally, several studies of the performance of fingerprint examiners have been performed. The most recent such study was published in May 2011. See Bradford T. Ulery et al., “Accuracy and Reliability of Forensic Latent Fingerprint Decisions,” 108 Proceedings of the National Academy of Sciences 7733 (May 10, 2011): see also Hearing Ex. 1, at 45 (citing five additional articles). The FBI also conducts proficiency examinations of its examiners, which—even if taken under conditions that “do not accurately represent [those] encountered in the field”—are of some value in assessing the reliability of individual examiners. See United States v. Baines, 573 F.3d 979, 990 (10th Cir.2009).

The court recognizes that the NAS Report called for additional testing to determine the reliability of latent fingerprint analysis generally and of the ACE–V methodology in particular. See NAS Report at 143–44. The Report also questions the validity of the ACE–V method. See id. at 142. However, Daubert, Kumho, and Rule 702 do not require “absolute certainty,” Daubert v. Merrell Dow Pharmaceuticals, Inc., 43 F.3d 1311, 1316 (9th Cir.1995) (opinion on remand); instead, they ask whether a methodology is testable and has been tested. On this record, the court finds that latent fingerprint analysis can be tested and has been subject to at least a modest amount of testing—some of which, like the study published in May 2011, was apparently undertaken in direct response to the NAS's concerns. The court therefore concludes that this factor weighs in favor of admitting latent fingerprint evidence. See Baines, 573 F.3d at 990 (“[W]hile we must agree with defendant that this record does not show that the technique has been subject to testing that would meet all of the standards of science, it would be unrealistic in the extreme for us to ignore the countervailing evidence.”).

B. Peer Review and Publication

The second factor in Daubert and Kumho concerns whether a methodology has been subject to publication and peer review. Publications regarding latent fingerprint analysis are relatively few in number. The government introduced into evidence a list of roughly thirty publications touching on various aspects of the field. See Hearing Ex. 1, at 66–69. Love's moving papers also contain citations to a handful of other relevant publications. See Mot. at 19–21. Although limited in quantity, many of these publications appear to “address ... theoretical/foundational questions” in latent fingerprint analysis. Mitchell, 365 F.3d at 239. The articles are therefore relevant to the reliability inquiry.

Ruth stated that at least one of these articles was published in a peer-reviewed journal, and the government's brief cites three other examples of peer-reviewed work. See Opp'n at 17. Even assuming that all of the articles cited by Ruth and the parties are peer-reviewed, latent fingerprint analysis would have only a small fraction of the number of peer reviewed publications found in established sciences. Nonetheless, because there are a handful of publications that concern the reliability of latent fingerprint analysis, at least a few of which are peer-reviewed, this factor is either neutral or weighs slightly in favor of admissibility. See, e.g., Boyd, 576 F.3d at 946 (affirming the admission of evidence of a theory supported by fourteen publications, ten of which were peer-reviewed); United States v. Prime, 431 F.3d 1147, 1153–54 (9th Cir.2005) (admitting handwriting analysis evidence in part because of publications that were peer reviewed by other forensic scientists).

C. Error Rates

The third factor deals with the known or potential error rate of a methodology. Latent fingerprint examiners have sometimes stated that their analyses have a zero error rate. See, e.g., NAS Report at 143. Ruth did not so testify. Rather, she stated that the ACE–V methodology is unbiased and that the methodology itself introduces no random error. Errors occur, but those errors are human errors resulting from human implementation of the ACE–V process. See Hearing Ex. 1, at 47–52. Because human errors are nonsystematic, Ruth believes that there is no overall predictive error rate in latent fingerprint analysis.

Ruth's testimony does not mean that the ACE–V process is perfect, or even that it is necessarily the best possible process for identifying latent prints. See NAS Report at 143 (“The [ACE–V] method, and the performance of those who use it, are inextricably linked, and both involve multiple sources of error (e.g., errors in executing the process steps, as well as errors in human judgment).”). Nevertheless, all of the relevant evidence in the record before the court suggests that the ACE–V methodology results in very few false positives—which is to say, very few cases in which an examiner identifies a latent print as matching a known print even though the two prints were actually made by different individuals. Most significantly, the May 2011 study of the performance of 169 fingerprint examiners revealed a total of six false positives among 4,083 comparisons of non-matching fingerprints for “an overall false positive rate of 0.1%.” See Ulery et al., supra, at 7733, 7735. This false positive rate is marginally higher than the rates the FBI has previously estimated. See, e.g., Barnes, 573 F.3d at 990–91 (recounting testimony suggesting that the FBI's error rate was 1 in 11 million cases). That discrepancy may be partially due to the fact that identifications in the study were not subject to verification.\2/ At least in Ruth's estimation, the error rate in the study may also have been marginally higher because the prints used in the study were of relatively low quality.\3/ In any case, a false positive rate of 0.1% remains quite low.\4/

2. No two individuals made the same false positive conclusion. See Ulery et al., supra, at 7735.

3. The examiners who participated in the study generally said that the prints in the study were similar in quality to those they encounter in their work. See Ulery et al., supra, at 7734.

4. The false negative rate in the May 2011 study was somewhat higher. See Ulery et al., supra, at 7733, 7736. The rate of false negatives is, however, not relevant here insofar as Ruth's testimony will identify various fingerprints as belonging to Love or others. See Mitchell 365 F.3d at 239. The FBI also concluded that two fingerprints in this case do not match those of Love or of the three individuals who pled guilty to the bombing. However, Love's brief does not specifically challenge the ACE–V methodology on the basis of its false negative rate, perhaps because a higher rate of false negatives could result if examiners take special care not to misidentify prints. See id. at 239 n. 19. The rate of false negatives may therefore be inversely related to the rate of false positives.


To counter the evidence of a low false positive rate, Love relies heavily on a high-profile misidentification made by the FBI in its investigation into the terrorist bombing of a train in Madrid. However, one confirmed misidentification is in no way inconsistent with an exceedingly low rate of error. Nor is the only other evidence on which Love relies. An examiner who served for fourteen years on a board that investigates complaints of misidentification has stated that he knew of twenty-five to thirty misidentifications that occurred in the United States during those fourteen years. See Mot. at 57 (citing Office of the Inspector General, A Review of the FBI's Handling of the Brandon Mayfield Case 137 (2006), available at: ht tp:// www.justice.gov/oig/special/s0601/final.pdf). Of course, any misidentification is troublesome. Without more foundation, however, this statement does not translate into a quantifiable error rate.\5/ The Inspector General's report does not include a benchmark for comparison.

5. For example, some rudimentary math suggests that this statement may imply a very low error rate. Ruth testified that she has made 1,200 identifications in six years as an examiner. At the same rate, she would perform 2,800 identifications over fourteen years, and it would take only ten examiners to perform 28,000 identifications over that time—roughly the number needed for 25–30 false positives assuming a 0.1% false positive rate. Because there are undoubtedly many more than ten latent fingerprint examiners in the United States, the examiner's statement is, if anything, indicative of a false positive rate lower than 0.1%.

 The court acknowledges that, as Ruth testified, historical error rates do not necessarily reflect the predictive error rate that should be applied to identifications made by any given examiner. See also United States v. Cordoba, 194 F.3d 1053, 1059 (9th Cir.1999) (discussing polygraphs). However, there is no evidence in the record to suggest that the rate of misidentifications made by latent fingerprint examiners is greater than an average of 0.1%, and some evidence suggests that the rate is much lower than that figure. The court therefore concludes that the error rate favors admission of latent fingerprint evidence. Accord, e.g., United States v. John, 597 F.3d 263, 275 (5th Cir.2010); Mitchell, 365 F.3d at 241; see also United States v. Chischilly, 30 F.3d 1444, 1154–55 (9th Cir.1994) (“conclud[ing] that there was a sufficient showing of low error rates” in DNA profiling despite arguments that “the potential rate of error in the forensic DNA typing technique is unknown”) (internal quotation marks omitted).

D. Standards

The fourth factor discussed in Daubert and Kumho is whether standards control a technique's operation. It is not disputed that the ACE–V methodology leaves much room for subjective judgment. According to Ruth, a fingerprint examiner must decide whether a latent print has value for comparison. Then, if the print does have value, the examiner must choose both the points at which to begin her comparison of the latent print to the known print and the sequence of comparisons to make from those starting points. Examiners in the United States also make a subjective decision at the evaluation stage; it is up to the examiner to determine whether or not prints match on the basis of the comparison and her experience and expertise. The verifying examiner then makes the same subjective evaluation as the original examiner. Love suggests that the standards factor weighs against admission because these subjective elements in the ACE–V process imply a lack of specific, objective standards that guard against examiner bias and error.

The standards used to guide an examiner's judgment vary across laboratories. There is not a single national standard. In part for this reason, and in part because of the subjective judgments made during the ACE–V process, the court acknowledges that the standards used in fingerprint analysis “are insubstantial in comparison to the elaborate and exhaustively refined standards found in many scientific and technical disciplines.” Mitchell, 365 F.3d at 241.

The court finds, however, that various standards imposed by the FBI's latent fingerprints unit, which conducted the analyses relevant to this case, sufficiently safeguard against bias and error. The FBI uses three different types of standards in an effort to reach reliable conclusions through the ACE–V process. At the laboratory level, the FBI adheres to the standards for calibration laboratories provided by the International Organization for Standardization and the standards for forensic laboratories promulgated by the American Society of Crime Laboratory Directors Laboratory Accreditation Board (“ASCLD/LAB”). See Prime, 431 F.3d at 1153–54 (noting that a laboratory conformed to ASCLD/LAB standards). The FBI also follows the consensus-based guidelines for laboratories engaged in friction ridge analysis issued by the Scientific Working Group on Friction Ridge Analysis Study and Technology.

At the level of individual examiners, the FBI applies relatively stringent standards for qualification. In addition to a college degree, examiners must have a significant amount of training in the physical sciences. They are then given eighteen months of FBI training and a four-day proficiency examination. After passing that test, an examiner has a six-month probationary period in which every aspect of her work is fully reviewed. Even thereafter, examiners undergo annual audits, proficiency tests, and continuing education.

Finally, at the level of individual comparisons and evaluations, the ACE–V methodology provides procedural standards that must be followed, such as the requirement that examiners assess each ridge and ridge feature in the prints under comparison. See United States v. Crisp, 324 F.3d 261, 269 (4th Cir.2003). The FBI has also recently incorporated documentation requirements that record the process used to detect latent prints as well as the examiner's comparison of the latent print to a known print. These documentation requirements and other procedures are enforced through the technical and administrative review of examiners' work. And the verification process serves as an error-reducing backstop.\6/

6. Under certain relatively rare circumstances, the FBI now performs verifications that are blind to both the identity of the original examiner and the conclusion that examiner reached. Blind verification was used for only one print at issue in this case.

In short, despite the subjectivity of examiners' conclusions, the FBI laboratory imposes numerous standards designed to ensure that those conclusions are sound. See United States v. Llera–Plaza, 188 F.Supp.2d 549, 571 (E.D.Pa.2002) (concluding that the subjective elements in latent fingerprint analysis are of a significantly “restricted compass”). The court therefore concludes that this factor weighs in favor of admission.\7/

7. Love's brief points to an ongoing controversy regarding whether a minimum number of matching details should serve as a prerequisite to the identification of a latent fingerprint. Ruth testified that a minimum points standard would be misguided, because three matching details in an area of the finger containing relatively few characteristics can be more telling than, say, eight matching points in a detail-rich area. In holding that the standards factor supports admission, the court does not attempt to resolve this dispute. The court instead concludes that other standards unrelated to a minimum points standard provide sufficient guidance to satisfy Daubert, Kumho, and Rule 702.

E. General Acceptance

The final factor discussed in Daubert and Kumho is whether a technique is generally accepted in a relevant scientific community. Love argues that the NAS report's criticisms of latent fingerprint analysis in general and the ACE–V methodology in particular demonstrate that friction ridge analysis is not accepted in the relevant scientific community. That assertion contains a kernel of truth. The NAS report does demonstrate some hesitancy in accepting latent fingerprint analysis on the part of the broader scientific community. Love's claim is subject to two significant qualifications. First, the NAS report itself states that “friction ridge analysis has served as a valuable tool, both to identify the guilty and to exclude the innocent.” NAS Report at 142. Instead of a full-fledged attack on friction ridge analysis, the report is essentially a call for better documentation, more standards, and more research. Cf. Chischilly, 30 F.3d at 1154 (“[T]he mere existence of scientific institutions that would interpret data more conservatively scarcely indicates a ‘lack of general acceptance’ ....”).

Second, Love does not dispute that the forensic science and law enforcement communities strongly support the use of friction ridge analysis. Acceptance in that narrower community is also relevant to the Daubert inquiry. See, e.g., Baines, 573 F.3d at 991 (stating that “acceptance of other experts in the field should ... be considered”); Mitchell, 365 F.3d at 241 (“[W]e consider as one factor in the Daubert analysis whether fingerprint identification is generally accepted within the forensic identification community.”); see also Prime, 431 F.3d at 1154 (affirming the admission of handwriting evidence in part because the district court “recognized the broad acceptance of handwriting analysis and specifically its use by such law enforcement agencies as the CIA, FBI, and the United States Postal Inspection Service”). For both of these reasons, the court concludes that the general acceptance factor at least weakly supports the admission of latent fingerprint evidence.

F. Relationship to Established Techniques

“[A] a court assessing reliability may consider the ‘novelty’ of the new technique, that is, its relationship to more established modes of scientific analysis.” Downing, 753 F.2d at 1238. Insofar as such a relationship exists, “the scientific basis of the new technique [may have] been exposed to” indirect “scientific scrutiny.” See id. at 1239. The Third Circuit held in Mitchell, and Ruth testified, that friction ridge analysis is related to the undisputeclly scientific “fields of developmental embryology and anatomy,” which explain “the uniqueness and permanence of areas of friction ridge skin.” Mitchell 365 F.3d at 242. Love argues that this strong tie to the biological sciences is irrelevant, because “ ‘uniqueness and persistence are necessary” ‘ but not sufficient conditions for friction ridge analysis to be reliable. Mot. at 67 (quoting NAS Report at 144). Even if uniqueness and persistence alone cannot validate latent fingerprint analysis, however, it remains true that “[i]ndependent work in [developmental embryology and anatomy] bolsters” two critical “underlying premises of fingerprint identification.” Mitchell, 365 F.3d at 242. This factor weighs in favor of admission.

G. Non–Judicial Applications

“[N]on-judicial use of a technique can imply that third parties ... would vouch for the reliability of the expert's method.” Id. at 242–43. Ruth testified that friction ridge analysis is used for various other purposes, including to identify disaster victims, to identify newborns at hospitals, for biometric devices, for some passports and visas, and for certain jobs. See Hearing Ex. 1, at 64–65. Love stresses that these non-judicial applications generally use rolled and not latent prints. Ruth testified, however, that latent prints are used to identify disaster victims when rolled prints are unavailable. See also Barnes, 573 F.3d at 990 (noting the use of latent prints in this context). But see Mitchell, 365 F.3d at 243 (stating that post-disaster identifications “differ from latent fingerprint identification because [those] identification[s] us[e] actual skin[, which] eliminates the challenges introduced by distortions”). On the basis of the widespread use of fingerprints, and occasional use of latent prints, for non-judicial identification purposes, the court concludes that this factor modestly supports the admission of latent fingerprint evidence.

H. Summary

The court recognizes that the NAS Report and other publications cited by Love critique some aspects of latent fingerprint analysis. However, the forensic science community generally and the FBI in particular have begun to take appropriate steps to respond to that criticism. On this record, in part because of recent developments regarding testing, publication, error rates, and the FBI's governing standards, none of the seven factors discussed by the parties weighs against the admission of latent fingerprint evidence. Friction ridge analysis is not foolproof, but it is also far removed from the types of “junk science” that must be excluded under Rule 702, Daubert, and Kumho. Considering and weighing all of the factors, the credible testimony of Ruth, and the written submissions by both parties, the court concludes that latent fingerprint analysis is sufficiently reliable to be admitted. The court denies Love's motion to exclude the testimony of Robin Ruth insofar as that motion is based on the supposed unreliability of latent fingerprint analysis generally. See, e.g., John, 597 F.3d at 276 (affirming the admission of latent fingerprint evidence); United States v. Pena, 586 F.3d 105, 110–11 (1st Cir.2009) (same); Baines, 573 F.3d at 992 (same); Mitchell, 365 F.3d at 245–46 (same); Crisp, 324 F.3d at 269 (same); United States v. Hernandez, 299 F.3d 984, 991 (8th Cir.2002) (same); United States v. Havvard, 260 F.3d 597, 601 (7th Cir.2001) (same); see also United States v. Sherwood, 98 F.3d 402, 408 (9th Cir.1996) (holding that it was not error to admit fingerprint evidence when the party challenging the evidence admitted that several Daubert factors were satisfied).

III. The Evidence in this Case

Love's motion to exclude the specific evidence the government wishes to introduce in this case also implicates the questions of whether Robin Ruth is qualified to testify as an expert and whether Ruth's analyses are reliable.

There can be no doubt that Ruth is qualified to testify as an expert in latent fingerprint analysis. Since October 2010, Ruth has served as the standards and practices program manager in the FBI's latent prints unit, meaning that she is tasked with overseeing the quality control efforts in that unit. She previously spent five years as an FBI fingerprint examiner; since joining the FBI, she has conducted over 150,000 comparisons and has made roughly 1,200 fingerprint identifications. Ruth possesses degrees in genetic engineering and in forensic and biological anthropology, see Opp'n, Ex. A, at 2, has co-authored two articles about friction ridge analysis, and routinely teaches and lectures about the field, see id. at 4–5.

Love argues that Ruth should not be allowed to testify because she is not certified by the International Association for Identification (“IAI”).\8/ Love's contention relies on the NAS Report, but that report does not state that all examiners should be IAI-certified. It instead simply recommends that forensic scientists be accredited by some outside body, and suggests that the not-yet-existing National Institute of Forensic Science should “determin[e] appropriate standards for accreditation and certification.” See NAS Report 208, 215. Rule 702 requires only that an expert be “qualified as an expert by knowledge, skill, experience, training, or education.” “No specific credentials or qualifications are mentioned,” and the Ninth Circuit has “held that an expert need not have [any] official credentials in the relevant subject matter to meet Rule 702's requirements.” United States v. Smith, 520 F.3d 1097, 1105 (9th Cir.2008). The court rejects Love's challenge to Ruth's qualifications.

8. Ruth has never taken the IAI's certification test for fingerprint examiners, but she is an IAI member, and she passed the FBI's own proficiency test.

Love does not argue that the evidence in this case is less reliable than other latent fingerprint evidence, and there is no reason to believe it is. The fingerprint evidence in this case was analyzed using the ACE–V methodology and the FBI's standard operating procedures. See Sherwood, 98 F.3d at 408 (noting that the examiner's “technique [was] the generally-accepted technique for testing fingerprints”). Ruth is the actually the third FBI examiner to analyze the evidence at issue: After an examiner performed the initial comparison and evaluation, those steps were verified by the senior examiner who performs technical reviews for the FBI's latent print unit. Ruth then re-examined the prints. All three examiners reached the same conclusions regarding each of the fifteen prints at issue. Ruth also testified that no recourse to Level 3 details—the very small details in a print like pores, which Love contends are easily misinterpreted—was necessary to reach or support her conclusions. In short, nothing about the latent prints in this case suggests that Ruth's conclusions are less reliable than other conclusions reached using the ACE–V method as implemented by the FBI. Those conclusions are therefore sufficiently reliable to be admitted into evidence.\9/ Of course, Ruth will be subject to cross-examination about her background, methods, analysis, conclusions, and latent fingerprint analysis generally.

9. It is undisputed that the fingerprint evidence—which includes evaluations of latent prints taken from literature related to the construction of pipe bombs—is highly relevant to this case. 

For these reasons, it is ORDERED that Love's motion to exclude the testimony of Robin Ruth is DENIED.

Friday, 15 June 2012

CODIS Loci Ready for Disease Prediction, Vermont Court Says

A trial court in Vermont has gone where no court has gone before. In State v. Abernathy [1], Chittenden Superior Court Judge Alison Sheppard Arms found that because "[s]ix CODIS loci ... have associations with an increased risk of disease or have functional properties," the custodians of law enforcement DNA databases can make "probabilistic predictions of disease." According to the judge, modern research has established that "some of the CODIS loci have associations with identifiable serious medical conditions," making the scientific evidence "sufficient to overcome the previously held belief[s]" about the innocuous nature of the CODIS loci.

Emphasizing this finding that richly information-laden STR profiles reside in identification databases, the court proceeded to strike down "Vermont's new pre-conviction DNA testing requirement ... that requires submission of a DNA sample from a 'person for whom the court has determined at arraignment there is probable cause that the person has committed a felony ... .'" In an atypical opinion, the court applied a "special needs" balancing test, placed the burden of proof on the state, and held that this law violates the state constitution.

A major theme in Judge Arms' discussion of human genetics is that there has been a revolution in our understanding of what used to be called "junk DNA." Even though the CODIS loci originally were described as "junk" in "good faith," that understanding was wrong--we now know that even DNA that does not code for proteins is biologically important.1 Other judges, advocacy groups, and at least one law professor have jumped from the discovery that the triplet code for proteins is not the sole message inscribed in DNA to the conclusion that all the CODIS loci may well convey significant information about disease states or propensities.

There are a couple of problems with this reasoning. All that we actually know is that some non-protein-coding DNA regulates gene expression. Scientists do not believe that all non-protein-coding sequences are regulatory. In particular, whether noncoding, nontranscribed, and largely nonconserved sequences are part of a regulatory system (even if their presence might have some function) is far from established.2 The opinion cites an essay I wrote making this point [3] but then ignores its content. It quotes the legal treatise, Modern Scientific Evidence, for the view that "while it is generally agreed that no single loci [sic] contains a gene that definitively determines any discernible characteristic of significance, there are nonetheless indications that they may play a role in some sensitive matters, and continued debates about their importance." Before Abernathy, it appeared that the "continued debates" ended five years ago with agreement with what already was clear -- that even if the loci do not play a functional role, they might, like certain fingerprint patterns or blood types, have some statistical associations with diseases.3

Venturing beyond the inconclusive generalities like these, Abernathy refers to the biomedical literature on five loci and to a testifying expert's characterization of the literature (with no specific references) on another locus. The opinion does not give the magnitude of any putative association, let alone any measure of predictive utility.4 It uses the following phrases: "a fairly large effect size," "a modest association," "not the most strongly associated," "small but ... not zero,"5 and "cannot find that this marker has no association." It does not provide measures of the uncertainty in these estimates. Finally, the opinion does not discuss the extent to which the studies said to show that the associations are real have been replicated.6

Of course, few judges could confidently review the flood of studies on human genetics. Unlike some previous opinions and law review articles, however, this opinion does not rely entirely or largely on newspaper headlines and stories about "junk DNA." Here, the iconoclastic findings came after an evidentiary hearing. But, as has happened before with DNA evidence [8], the evidentiary hearing was one-sided. The defendants presented the testimony of Professor Gregory Wray of Duke University, a specialist in genetics and evolutionary biology, and the state chose not to present an expert in medical genetics or genomics to counter his testimony. Although Professor Wray reviewed the biomedical literature before he testified, the defense submitted no written report, and the state rather than the defense introduced the papers cited in the opinion as exhibits. Scanning the testimony, it seems to me that Dr. Wray never was asked a series of critical questions:
  1. Is it generally accepted that the associations he pointed to apply to the population of individuals whose DNA is placed in law enforcement databanks?
  2. Assuming that they do apply to that population, what is the positive and negative predictive value of any inferences about disease based on the CODIS alleles?
  3. How would the predictive or diagnostic disease-related information in a state DNA database compare to that of (a) color photographs, (b) fingerprints, (c) the blood types used in old-fashioned serology, and (d) the HLA-A and HLA-B haplotypes once used on parentage testing?
  4. Are the CODIS genotypes likely to be substantially more predictive in the future?
Until these questions are answered, there is reason to ask whether the trial court's findings fairly represent the scientific status quo or instead are grim predictions of what could come to pass.

Notes

1. For a short audio clip reporting on the revolutionary discoveries, click on Joe Palca, Don't Throw It Out: 'Junk DNA' Essential In Evolution, All Things Considered, Aug. 19, 2011 (with a sound bite from Professor Gregory Wray, among other interviewees).

2. According to Judge Arms,"[t]he term 'junk DNA' was coined in the early 1980s." In fact, the phrase normally is attributed to Susumu Ohno, who used it in the title of a 1972 paper [2]. Ohno did not reason that "we don't know what noncoding DNA does, therefore, is it is useless junk." Indeed, he proposed that the duplication and inactivation of genes produce non-protein-coding DNA (now designated pseudogenes) that might have a function. A video introducing Ohno and reading an excerpt from the paper about the role of the noncoding sequences as "spacers" with evolutionary importance can be found at http://www.youtube.com/watch?v=nomI35DJB40&noredirect=1. Since 1972, other possible functions for noncoding DNA have been proposed. Some functions imply that the sequences should be conserved as one species evolves into another. Others, such as Ohno's suggestion that noncoding sequences act as buffers between genes, do not.

3. See [5, p. 228] (referring to "a brief debate in the legal literature" necessitated by "a misunderstanding by Simon Cole over some of the things I [John Butler] had written in a review article on STR markers" and emphasizing that "STR markers used for human identity testing do not predict disease."). One source of confusion, which also infects the Abernathy opinion is the thought that a statistical association between a locus and a disease detected in a family study in say, Northern India, establishes that the same association exists throughout the United States population.

4. Even a strong association (large relative risk) would not make for a useful predictive test if the prevalence of the condition is very small. See [3].

5. The sentence "[t]he relative risk of developing schizophrenia associated with this marker is small but it is not zero" is technically flawed. A relative risk of 1 would express a 0 correlation.

6. Replication is always important, and the problem of false positives is especially acute with genome-wide association studies. See, e.g., [6, 7].

References

1. State v. Abernathy, No. 3599-9-11 (Vt. Super. Ct. June 1, 2012).

2. S. Ohno, So Much "Junk" DNA in our Genome, 23 Brookhaven Symp. Biol. 366 (1972) (also published in Evolution of Genetic Systems 366 (H.H. Smith ed. 1972).

3. David H. Kaye, Please, Let's Bury the Junk: The CODIS Loci and the Revelation of Private Information, 102 Nw. U. L. Rev. Colloquy 70 (2007).

4. David H. Kaye, Mopping Up After Coming Clean About "Junk DNA", Nov. 23, 2007, available at http://ssrn.com/abstract=1032094.

5. John M. Butler, Advanced Topics in Forensic DNA Typing: Methodology (2012).

6. D.J. Hunter & P. Kraft, Drinking from the Fire Hose--Statistical Issues in Genomewide Association Studies, 357 N. Engl. J. Med. 436 (2007).

7. Thomas A. Pearson, & Teri A. Manolio, How to Interpret a Genome-wide Association Study, 299 J. Am. Med. Ass'n 1335 (2008).

8. David H. Kaye, The Double Helix and the Law of Evidence (2010).

Cross-posted from The Double Helix Law Blog.