Showing posts with label error. Show all posts
Showing posts with label error. Show all posts

Friday, 6 June 2014

Fingerprinting Errors and a Scandal in St. Paul

Reports of crime laboratory scandals and fraud have become legion. It is scandalous to convict and imprison people on the basis of contrived, fabricated, or even incompetently generated laboratory evidence. But not all scandals are equal. Consider, for example, the report last year in the ABA Journal that
the St. Paul, Minn., police department’s crime lab suspended its drug analysis and fingerprint examination operations after two assistant public defenders raised serious concerns about the reliability of its testing practices. A subsequent review by two independent consultants identified major flaws in nearly every aspect of the lab’s operation, including dirty equipment, a lack of standard operating procedures, faulty testing techniques, illegible reports, and a woeful ignorance of basic scientific principles. 1/
The article proceeds to describe deplorable conditions in the drug testing lab, but all it says about latent print work is that “[t]he city has since hired a certified fingerprint examiner to run the lab, who has announced plans to resume its fingerprint examination and crime scene processing operations, and begin the procedure for seeking accreditation.”

Curious as to what the latent-print examiners had been doing, I turned to a local newspaper article entititled "St. Paul Crime Lab Errors Rampant." It reported that “[t]he police department hired two consultants to work on improving the lab after a ... Court hearing last year disclosed flawed drug-testing practices” and that “the lab recently resumed fingerprint work by certified analysts.” 2/

One consultant, “Schwarz Forensic Enterprises of Ankeny, Iowa, ... studied the crime lab's latent fingerprint comparison, processing and crime scene units” and found that “[p]ersonnel appeared to have attended seminars and training, but there wasn't formal competency testing or a program to assess ongoing proficiency.”

These untested personnel offered an opportunity to see how poorly monitored analysts performed. Would they succumb to the widely advertised cognitive biases that might cause latent print examiners to declare matches that do not exist? Would they declare matches more frequently than certified examiners? Apparently not:
"'Despite these deficiencies, no evidence of erroneous identifications by latent print examiners was found; but we did find numerous examples of cases wherein examiners had failed to claim latent prints as suitable for identification and/or to identify prints to suspects,'
the Schwarz report said." In other words, the incidence of false negatives and missed opportunities to make identifications or exclusions was high, but no false-positive errors were found. “A review of 246 fingerprint cases found the unit successfully identified prints only ‘in cases where the print detail is of extraordinarily high quality.’”

This outcome is consistent with more rigorous studies showing that when latent print examiners make mistaken comparisons, the errors are usually false exclusions—not false matches. 3/ This tendency reflects a different sort of bias—an unwillingness to declare a match unless the match seems quite clear.

Of course, 246 instances without false positives from worrisome fingerprint analysts does not prove that they never make false matches. If this group were making false identifications 1% of the time, for instance, the probability that no false positives would be seen in a run of 246 independent cases (each with the same 1% false-match probability) would be (1 – .01)246 = 8%.

The absence of false positives also is consistent with an intriguing 2006 report by Itiel Dror and David Charlton. 4/ These investigators had six experienced, certified, and proficiency-tested analysts examine sets of prints from four cases in which, years ago, the examiners had found exclusions and another four cases in which they had made identifications. The subjects did not realize that they had seen these prints before. In some instances of previous exclusions, the examiners were told that a suspect had confessed. In none of these cases did the examiners depart from their earlier judgment of a match.

On the other hand, in cases of previous identifications, when examiners were told that the suspect was in police custody at the time of the crime, two examiners switched from an exclusion to an identification, and one switched to “cannot decide.” Although these sample sizes are too small to justify strong and widely generalizable conclusions, it looks like it is easier for information that is not needed for the analysis to prompt an exclusion than an individualization.

Dror and Charlton interpret their results as supporting (among other things) the claim “that the threshold to make a decision of exclusion is lower than that to make a decision of individualization.” This higher threshold would make it more difficult to bias an examiner to make a false identification than to make a false exclusion.

Did any of the 246 St. Paul cases involve contextual bias of one kind or another? If so, it would be interesting to find out if these examiners resisted contextual suggestions favoring identifications or exclusions in those cases. Audits like these could be helpful not only in getting laboratories with problems back on track, as in St. Paul, but also as a source of information on the risks of different types of errors in various settings and circumstances.

Notes
  1. Mark Hansen, Crime Labs Under the Microscope after a String of Shoddy, Suspect and Fraudulent Results, ABAJ, Sept. 2013
  2. Mara H. Gottfried & Emily Gurnon, St. Paul Crime Lab Errors Rampant, Reviews Find, Pioneer Press, Feb. 14, 2013
  3. See, e.g., Fingerprinting Under the Microscope: Error Rates and Predictive Value, Forensic Science, Statistics, and the Law, April 30, 2012; Fingerprinting Error Rates Down Under, June 24, 2012, Forensic Science, Statistics, and the Law.
  4. Itiel E. Dror & David Charlton, Why Experts Make Errors, 56 J. Forensic Identification 600-16 (2006)

Wednesday, 4 June 2014

Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 3)

After explaining that Florida’s statutory cutoff of –2σx corresponds to an IQ score of 70 (because IQ tests are normed to have a mean of 100 and a standard deviation of 15), Justice Kennedy observes that:
Florida's rule disregards established medical practice in two interrelated ways. It takes an IQ score as final and conclusive evidence of a defendant's intellectual capacity, when experts in the field would consider other evidence. It also relies on a purportedly scientific measurement of the defendant's abilities, his IQ score, while refusing to recognize that the score is, on its own terms, imprecise.
Here, I show that these two limitations on IQ scores are less “interrelated” than Justice Kennedy suggests.

The first issue: validity

The first limitation involves what social scientists call “validity”—the extent to which something measures the real quantity of interest. For example, measuring the volume of a box by attending only to one dimension, such as height, is invalid because it ignores the two other determinative variables of width and depth.

The second limitation concerns the precision or “reliabilility” of the measurement—regardless of validity. Being able to measure the height of a box to the nearest millimeter time after time achieves precision and reliability, but it still lacks validity (with respect to the variable of volume). Moreover, measuring width or depth—even somewhat imprecisely—will add to validity but will do nothing to enhance the precision of the measurement of height.

Likewise, the “other evidence” to which the Court referred does not make IQ scores any more precise. Rather, it relates to what is known in the trade as “adaptive functioning.” The opinion defines adaptive functioning as “the inability to learn basic skills and adjust behavior to changing circumstances.” The Court disparages the “mandatory cutoff” of –2σx because this cut-score means that
sentencing courts cannot consider even substantial and weighty evidence of intellectual disability as measured and made manifest by the defendant's failure or inability to adapt to his social and cultural environment, including medical histories, behavioral records, school tests and reports, and testimony regarding past behavior and family circumstances. This is so even though the medical community accepts that all of this evidence can be probative of intellectual disability, including for individuals who have an IQ test score above 70.
If one were to follow this we-need-another-variable theory of “intellectual disability” to its logical limit, no IQ score could preclude the more comprehensive assay of all forms of “substantial and weighty evidence of intellectual disability.” An individual with above-average IQ scores also might “manifest [a] failure or inability to adapt to his social and cultural environment [as shown by] medical histories, behavioral records, school tests and reports, and testimony regarding past behavior and family circumstances.”

But surely the Court cannot claim that the execution of intellectually gifted but maladapted criminals is cruel and unusual while the execution of intellectually gifted and socially well adjusted criminals is not. To avoid such anomalies, the Court follows the contemporary (and prior) mental health practice of limiting “intellectual disability” to “concurrent deficits in intellectual and adaptive functioning” (emphasis added), which requires “significantly subaverage intellectual functioning” in addition to mere “deficits in adaptive functioning.” If it is clear that an individual is able to function intellectually within a broad but “normal” range, then a state need not entertain a claim of “intellectual disability” based solely on problems in adaptive functioning. Therefore, an IQ within the normal range should suffice to displace the offender from the potentially death-disqualified group.

So the possibility of weighty evidence of deficits in adaptive functioning, although relevant to clinicians, turns out to be no explanation for why Florida cannot draw the line at –2σx. If evidence of an adaptive-function deficit does not bar the state from executing criminals with IQs in the broad range of normalcy, why cannot the state define all IQs above 70 (–2σx) as lying within that range? That “experts in the field would consider other evidence” than IQ scores is not an answer. The answer has to be that (1) there is a range above which IQ, in of of itself, is a valid measure of the absence of “intellectual disability,” and (2) this range does not extend all the way down to 70. When these conditions hold, experts would not (or need not) consider “other evidence.”

Ironically, the Court’s opinion contradicts the second proposition. It clearly implies that the state could use a perfectly precise IQ measurement just above –2σx as conclusive evidence of intellectual disability. But if that is so, then the problem is not the failure to allow evidence of adaptive functioning. It is solely the existence of nonzero measurement error of IQ alone.

The second issue: precision (reliability)

Apparently (and dubiously) reserving the term “scientific” for precise measurements, Justice Kennedy stated that the “purportedly scientific measurement of the defendant's abilities, his IQ score, ... is, on its own terms, imprecise.” The problem is that although “there is evidence that Florida's Legislature intended to include the measurement error in the calculation ... the Florida Supreme Court ... has held that a person whose test score is above 70, including a score within the margin for measurement error, does not have an intellectual disability ... .”

In other words, a legislature that wants to preclude the more elaborate evaluations of all offenders with IQ scores below 70 could do so if only it had a way to measure IQs with perfect accuracy. Because of the “measurement error” of IQ tests, this legislature must adopt a higher cutoff. The Court, relying on the diagnostic literature, repeatedly refers to a cutoff of 75 as assuring an adequate safety margin.

The dissent had harsh words for the choice of 75, and I will get to those later, after examining where the figure of 75 comes from. At this point, no excursion into statistical theory is required to recognize that there is something weird about saying that IQ scores are problematic because they are an incomplete measure of “intellectual disability,” but then using them—and only them—within a band that accounts only for the error in measuring IQ. By definition, this band does not attend to the other factors that should be part of the full analysis. To put it another way, if the problem lies with using IQ alone, the solution lies in defining the range of IQ scores in which the other factors realistically could produce a different diagnosis. However, the error in IQ measurements has no clear connection to the range in which the failure to look beyond IQ makes a difference.

The majority’s response is essentially that if the mental health profession generally agrees that incompleteness is only a significant concern within the logically unrelated range of IQ-score error, then that is all that the Cruel and Unusual Punishment Clause demands. To which the dissent replies that abdicating the line drawing to the professionals makes no constitutional sense and “will also lead to serious practical problems.”

The dissent’s peculiar proof of “instability”

The first such problem is “instability.” According to Justice Alito:
This danger is dramatically illustrated by the most recent publication of the APA, on which the Court relies. This publication fundamentally alters the first prong of the longstanding, two-pronged definition of intellectual disability that was embraced by Atkins and has been adopted by most States. In this new publication, the APA discards “significantly subaverage intellectual functioning” as an element of the intellectual-disability test. Elevating the APA's current views to constitutional significance therefore throws into question the basic approach that Atkins approved and that most of the States have followed. 1/
The American Psychiatric Association’s latest version of its venerable Diagnostic and Statistical Manual of Mental Disorders—the DSM-5—“was published in May 2013 amid a storm of controversy and bitter criticism.” 2/ In general, critics maintain that “D.S.M.’s diagnostic categories lacked validity, that they were not ‘based on any objective measures,’ and that, ‘unlike our definitions of ischemic heart disease, lymphoma or AIDS,’ which are grounded in biology, they were nothing more than constructs put together by committees of experts.” 3/ Neither opinion even hints at such turmoil. The majority genuflects to clinical expertise and guidelines. The dissent raises no questions about validity and subjectivity, but objects to substituting “the standards of professional associations, which at best represent the views of a small professional elite” for “the standards of the American people.”

As for “instability,” the DSM-5 has brought a profusion of new or redefined disorders, but it does not radically change the definition of “intellectual disability” or dispense with the criterion of “significantly subaverage intellectual functioning.” It simply substitutes the word “intellectual ... deficit” for “significantly subaverage.” The diagnostic criteria have remained remarkably similar over the 19 years between the DSM-4 and the DSM-5.

The DSM-5 specifies that “[t]he first diagnostic criterion that “must be met” is “A. Deficits in intellectual functions ... confirmed by ... both clinical assessment and standardized intelligence testing.” If the intelligence testing does not demonstrate subaverage performance, it is hard to see how it could confirm the existence of a meaningful deficit. Moreover, the DSM-5 elaborates, making it plain that significantly subaverage IQ remains a sine qua non for the diagnosis:
The essential features ... are deficits in general mental abilities (Criterion A) and impairment in everyday adaptive functioning ... (Criterion B) [with o]nset is during the developmental period (Criterion C). The diagnosis of ... is based on both clinical assessment and standardized testing ... . Intellectual functioning is typically measured with ... tests of intelligence. Individuals with intellectual disability have scores of approximately two standard deviations or more below the population mean, including a margin for measurement error (generally +5 points). On tests with a standard deviation of 15 and a mean of 100, this involves a score of 65–75 (70 ± 5).
Compare this to the DSM-4 (or the DSM-4-TR cited by Justice Alito, which uses the same words):
The essential feature of Mental Retardation is significantly subaverage general intellectual functioning (Criterion A) that is accompanied by significant limitations in adaptive functioning ... (Criterion B) [with] onset ... before age 18 years (Criterion C). ... General intellectual functioning is defined by the intelligence quotient ... obtained by assessment with ... intelligence tests ... . Significantly subaverage intellectual functioning is defined as an IQ of about 70 or below (approximately 2 standard deviations below the mean). It should be noted that there is a measurement error of approximately 5 points in assessing IQ, although this may vary from instrument to instrument ... . Thus, it is possible to diagnose Mental Retardation in individuals with IQs between 70 and 75 who exhibit significant deficits in adaptive behavior. Conversely, Mental Retardation would not be diagnosed in an individual with an IQ lower than 70 if there are no significant deficits or impairments in adaptive functioning.
Thus, there are wording changes over the 19 years from 1994 to 2013, but Criterion A remains Criterion A, IQ tests remain critical to the diagnosis, and the range of test scores that lend themselves to the diagnosis is the same. The APA has changed the emphasis somewhat, and it has spelled out the constructs a little more (in words not quoted here). Nevertheless, to claim that the shift “dramatically illustrate[s a] fundamental[] alter[ation in] ... the longstanding ... definition of intellectual disability” seems, well, melodramatic.

State laws that rely on –2σx plus a margin of safety for measurement error, are compatible with Atkins, Hall, DSM-4, and DSM-5. Of course, whether this is a logically or functionally appropriate manner of defining “intellectual disability” for purposes of capital punishment is open to debate. Resolving this debate requires a more detailed and accurate understanding of the concept of measurement error than the Hall opinions provide.

Footnotes
  1. The second problem is that “changes adopted by professional associations are sometimes rescinded.” This problem is just a form of instability. The third problem is hypothetical (thus far) as it relates to intellectual disability determinations: “what if professional organizations disagree? The Court provides no guidance for deciding which organizations' views should govern.” The fourth and final “practical problem” is actually conceptual—and quite important. “[D]efinitions of intellectual disability ... are promulgated for use in making a variety of decisions that are quite different from the decision whether the imposition of a death sentence in a particular case would serve a valid penological end. ... [I]n determining eligibility for social services, adaptive functioning may be much more important.”
  2. Nat’l Health Service Choices, Controversy over DSM-5: New Mental Health Guide, Aug. 15, 2013.
  3. Gary Greenberg, The Rats of N.I.M.H., New Yorker, May 16, 2013 (quoting Thomas Insel, the director of the National Institute of Mental Health). 

Other postings in this series
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 1) (introduction)
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2) (on standard deviation)

Monday, 2 June 2014

Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2)

Justice Kennedy described the Florida law that prompted the trial court to reject Hall’s claim of intellectual disability as follows:
Florida's statute defines intellectual disability for purposes of an Atkins proceeding as “significantly subaverage general intellectual functioning existing concurrently with deficits in adaptive behavior and manifested during the period from conception to age 18.” Fla. Stat. § 921.137(1) (2013). The statute further defines “significantly subaverage general intellectual functioning” as “performance that is two or more standard deviations from the mean score on a standardized intelligence test.” Ibid. The mean IQ test score is 100. The concept of standard deviation describes how scores are dispersed in a population. Standard deviation is distinct from standard error of measurement, a concept which describes the reliability of a test and is discussed further below. The standard deviation on an IQ test is approximately 15 points, and so two standard deviations is approximately 30 points. Thus a test taker who performs “two or more standard deviations from the mean” will score approximately 30 points below the mean on an IQ test, i.e., a score of approximately 70 points.
Because standard deviations are fundamental to the Florida law, to the Court’s conclusions about it, and to its dicta regarding the lowest mandatory IQ cut-off that a state can use, I am going to be persnickety in unpacking this paragraph.

Although the Court is only discussing the standard deviation of IQ scores, “standard deviation” (SD) has a much broader meaning. When one considers the general meaning of the term, it becomes clear that the SD of the scores does not quite describe “how scores are dispersed in a population.” It merely indicates how much they are dispersed in either a population or a sample.

For example, the trial court heard testimony about at least four IQ test scores for Hall—71, 72, 73, and 80. (It declined to consider a score of 69 on another test because the psychologist who administered and scored the test was dead and Hall’s counsel had violated an order “to provide the State with the [underlying] testing materials and raw data.” Indeed, Hall had taken as many as nine IQ tests over a 40-year period.)  The standard deviation of the scores considered by the trial court is the square root of the average squared deviation from the mean—namely,

SD = {[(71–74)2 + (72–74)2 + (73–74)2 + (80–74)2]/4}1/2 = 3.53.

Tossing in the excluded score of 69 increases the SD to 3.74. The SD increases because the additional score is below the range of the other four, thus creating more variability in the sample (and lowering the mean from 74 to 73).

Of course, the Court’s number of 15 for the SD of IQ scores does not come from Hall’s scores. At this point, I use his scores only to elaborate on the Court’s observation that a standard deviation is a statistic that indicates how much the numbers in some set of numbers fluctuate around their mean. The standard deviation of 15 IQ points is an estimate of how much the scores of everyone in the general population—a large batch of numbers indeed—would vary if everyone took the test. The average score would be approximately 100, and there would be a lot of scatter around this mean. (In fact, the raw scores on the test are transformed in light of their mean and SD to force them to have a desired mean and SD near 100 and 15, respectively.) And, yes, 100 – (2×15) = 70, so Florida’s choice of 2 SDs to demarcate “significantly subaverage general intellectual functioning” translates into a score of 70 on a test with this mean and SD.

But the fact that every batch of numbers has a SD does not tell us “how [these numbers] are dispersed.” The numbers could be highly concentrated around a single value, with outliers on the flanks. Their distribution could be flat, with an equal fraction of the numbers spread out everywhere. The distribution might show clustering at several locations, and so on.

IQ scores, however, are dispersed approximately according to a “normal” or “Gaussian” curve. This distribution is the bell-shaped one prominent in elementary statistics courses. There are other bell-shaped curves, and all kinds of other interesting and important families of curves, but IQ scores, like many physical variables (such as weight and height), tend to be normally distributed across the members of a population (and hence in representative samples of that population).

The exact shape of all such normal distributions can be determined from two numbers—the mean and the standard deviation. The mean states where the bell sits, and the standard deviation determines how steeply its sides flow down from the top.You can see for yourself by entering your favorite means and standard deviations into the demonstration program in the OnlineStatBook.

Using the variable X to denote IQ scores and the symbol σx to designate their standard deviation, the particular normal distribution used in the Court’s calculation is such that, 2.28% of the scores lie below 70 (which, as the Court calculated it, corresponds to –2σx), and 4.75% fall below 75 (which, for the mean of 100 and standard deviation σx of 15, corresponds to –1.67σx). The latter IQ score, x = 75, is significant because Hall conceded (and the Court seemed to agree) that Florida could have chosen this score as its cut-off. For example, the Court expressed dissatisfaction that, in light of its calculations, the effect of Florida’s cut-off of –2σx was to preclude legally effective “professional[] diagnose[s of] intellectual disability [in a case like Hall’s, for which] the individual's IQ score is 75 or below.”

The dissent insisted that states should have more discretion to set cut-off scores. Unless –1.67σx (or 75) corresponds to the level of impairment that justifies a categorical rule, the majority has no satisfying reason to select one cut-off over the other. Why is the Court’s choice of 1.67 standard deviations below the mean the highest that the Constitution permits? Why is Florida’s two-standard-deviation rule insufficient?

The Court’s answer leans heavily on the standard error of measurement — another technical term that appears in the paragraph quoted above: “Standard deviation is distinct from standard error of measurement, a concept which describes the reliability of a test and is discussed further below.” But the standard error of measurement is also a standard deviation, one that is estimated, almost magically, from test reliability statistics. Thus, a more precise sentence would have been: “The standard deviation of all test scores is distinct from another standard deviation known as the standard error of measurement, which depends on the reliability of the test. We discuss the standard error of measurement below.”

Other postings in this series

  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 1), May 29, 2014 (introduction)
  •  Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2), June 2, 2014 (on standard deviation)
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 3), June 4, 2014 (on validity and the stability of the APA's diagnostic criteria)

Saturday, 31 May 2014

Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 1)

In Hall v. Florida, 134 S.Ct. 1986 (2014), the Supreme Court struck down Florida’s practice of using an IQ score of 70 points or less as a dispositive measure of the level of the “intellectual disability” that precludes capital punishment. Justice Kennedy’s opinion (joined by Justices Breyer, Ginsburg, Sotomayor, and Kagan) summarizes the case pithily:
This Court has held that the Eighth and Fourteenth Amendments to the Constitution forbid the execution of persons with intellectual disability. Atkins v. Virginia, 536 U.S. 304, 321 (2002). Florida law defines intellectual disability to require an IQ test score of 70 or less. If, from test scores, a prisoner is deemed to have an IQ above 70, all further exploration of intellectual disability is foreclosed. This rigid rule, the Court now holds, creates an unacceptable risk that persons with intellectual disability will be executed, and thus is unconstitutional.
Id. at 1990. Led by Justice Alito, Chief Justice Roberts and Justices Scalia, Kennedy, and Thomas dissented. The Court, they maintained, was overruling Atkins and adopting “a uniform national rule that is both conceptually unsound and likely to result in confusion.” Id. at 2002 (Alito, J., dissenting). Among other things, the dissent warns that the Court “misunderstands” the statistical concepts of standard error and confidence intervals, id. at 2009, and that it therefore “makes factual mistakes that will surely confuse States attempting to comply with its opinion.” Id. at 2010.

The dissent has a point. Parts of the majority opinion are elliptical and potentially confusing. Nonetheless, some of the harsh critique is overdrawn. Moreover, Justice Alito's presentation of psychometric concepts also is hardly error-free, inviting a rejoinder of "tu quoque."

Therefore, in a series of postings yet to come, I will, in Justice Alito's words, "wade[] into technical matters that must be understood in order to see where the Court goes wrong." But I'll do the same for the dissenting opinion's presentation of these matters.

Other postings in this series
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2), June 2, 2014 (on standard deviation)
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 3), June 4, 2014 (on validity and the stability of the APA's diagnostic criteria)

Sunday, 26 January 2014

Hundreds of Errors in DNA Databases: What Do They Mean?

The other day, the New York Times reported that "[t]he Federal Bureau of Investigation, in a review of a national DNA database, has identified nearly 170 profiles that probably contain errors" and that "New York State authorities have turned up mistakes in DNA profiles in New York’s database."

Apparently, nearly all the errors involve a recorded profile that differs from the true profile at one (and only one) allele. These "mistakes were discovered in July, when the F.B.I., using improved software, broadened the search parameters used to detect matches. The change, one F.B.I. scientist said, was like upgrading or refining 'a spell-check.' In 166 instances, the new search found DNA profiles in the database that were almost identical but conflicted at a single point."

This discovery raises several questions. How prevalent are these errors? What caused them? And, what investigative or prosecutorial errors could they cause?

I. Prevalence and Causes

The article observes that "[t]he errors identified so far implicate only a tiny fraction of the total DNA profiles in the national database, which holds nearly 13 million profiles, more than 12 million from convicts and suspects, and an additional 527,000 from crime scenes." Thus, "Alice R. Isenberg, the chief of the biometric analysis section of the F.B.I. Laboratory, said that ... 'We were pleasantly surprised it was only 166. ... These are incredibly small numbers for the size of the database.'" She believes "most of the 166 cases probably resulted from interpretation errors by DNA analysts or typographical errors introduced when a lab worker uploaded the series of numbers denoting a person’s DNA profile."

II. Consequences: Risks of False Hits and Misses

It is true that 166 is a small fraction--about 0.0013%--of all the profiles on record. But it would be interesting to know whether these discrepancies are concentrated in the known offender or arrestee profiles or in the ones from crime scenes (the "forensic index").

A. Errors in the Offender-arresstee Indices

Suppose first that the profile of an individual in an offender or arrestee index departs from the true profile of that individual by a single allele. If that individual committed an offense for which a crime-scene profile is recovered, he would not be noticed in a database trawl (unless the search program flagged near misses like this). This would be a false negative error.

How probable is this false negative error? There have been less than 230,000 hits (see http://www.fbi.gov/about-us/lab/biometric-analysis/codis/ndis-statistics) between profiles in the forensic index and the more than 12,000,000 profiles in the offender and arrestee indices. This means that an individual has about a 1.9% chance of being linked to a crime through the database. Assuming that all 166 misrecorded or mistyped profiles are in the offender-arrestee indices, it follows that the probability that one or more of these errors would yield a false negative is 166 x 0.0013% x 1.9% = 0.000042 = 0.0042%. Even if the 166 errors were the tip of the proverbial iceberg, 90% of which lies below the visible surface, the probability of a false negative resulting from the inaccurate profiles is only 0.042%.

If the inaccurate profile pertains to a offender or an arrestee, as we have assumed so far, then the probability of a false positive -- a hit to an individual represented in the database who is not the source of the crime-scene sample -- typically is even smaller. A false positive could occur if someone else in the world has a true profile that (1) is not in the offender-arrestee index and (2) perfectly matches the inaccurate profile in the index. Because all DNA profiles consisting of substantial number of loci are rare, the probability of a false positive must be small. 1/

B. Errors in the Forensic Index

Of course, all the inaccurate profiles did not come from previous offenders or arrestees. Some came from the crime-scene samples -- producing erroneous entries in the forensic index. These errors "had the effect of obscuring clues, blinding investigators to connections among crime scenes and known offenders" in three cases in New York. When forensic-index errors are present, the risk of a false negative error is much larger. Even if the source of the crime-scene sample is represented in the offender-arrestee indices (as seems to occur for some 2% of these database inhabitants), a database trawl for a match to the erroneous forensic index profile will miss this person.

Again, however, the chance of a false positive -- a hit between the false crime-scene profile and an offender or arrestee profile -- remains low for full DNA profiles. The probability of another person having the same false profile is still quite small. 2/

III. Caveats


The probabilities noted here are the result of back-of-the-envelope calculations (see note 1). Although I would think that more precise analyses would not give dramatically different results, I am not suggesting that the sources of errors reported by the FBI should be ignored or minimized. The incidence of errors in generating and recording data can be reduced, in part by automated systems. 3/ In addition, the estimates I have provided do not pertain to errors resulting from contamination of a crime-scene sample with an innocent suspect's sample and to profiles that are less complete than thirteen loci.

Notes
  1. If each STR allele is present in 10% of the relevant population (the actual values for each allele will vary from case to case) and if 13 loci are in the profile (the current norm), then the probability of a full match to the source of a crime-scene sample is less than (2 x 1/10)13 = 0.00000000082 (if the source of the crime-scene sample and the database inhabitant are unrelated members of a population in Hardy-Weinberg equilibrium and there is no linkage disequilibrium). Even if the correctly profiled crime-scene sample comes from a brother, the probability of a 13-locus match is only about (1/4 + p/2 + 2p2)13, where p is the average chance that the two alleles will match (the proportion of homozygotes in the population). (National Research Council Committee on DNA Technology in Forensic Science 1992, p. 167). If the homozygosity rate is, say 30% (which is higher than that reported in Budowle et al. (1999, p. 1278 (tbl. 1)), the probability of a full sibling possessing a one-off profile would only be about 0.00084.
  2. See supra note 1.
  3. Additional recording errors might be detected by massive, all-pairs trawls of the indices in CODIS. These would flag suspiciously similar profiles recorded as coming from what are thought to be different sources.

References

Saturday, 18 January 2014

The Signal, the Noise, and the Errors

Published in September 2012, The Signal and the Noise by Nate Silver soon reached The New York Times best-seller list for nonfiction. Amazon.com named it the best nonfiction book of 2012, and it won the 2013 Phi Beta Kappa Award in science. Not bad for a book that presents Bayes' rule as prescription for better thinking in life and as a model for consensus formation in science. The subtitle, for those who have not read it, is "Why So Many Predictions Fail--But Some Don't," and the explanations include poor data, cognitive biases, and statistical models not grounded in an understanding of the phenomena being modeled.

The book is both thoughtful and entertaining, covering many fields. I learned something about meteorology (your local TV weather forecaster probably is biased toward forecasting bad weather--stick to the National Weather Service forecasts), earthquake predictions, climatology, poker, human and computer chess-playing, equity markets, sports betting, political polling and poor prognostication by pundits, and more. Silver does not pretend to be an expert in all these fields, but he is perceptive and interviewed a lot of interesting people.

Indeed, although Wikipedia describes Silver as "an American statistician and writer who analyzes baseball (see Sabermetrics) and elections (see Psephology)," he does not present himself as as expert in statistics, and statisticians seem conflicted on whether to include him their ranks (see AmStat News). He seems to be pretty much self-educated in the subject, and he advocates "getting your hands dirty with the data set" rather than "spending too much time doing reading and so forth." Frick (2013).

Perhaps that emphasis, combined with the objective of writing an entertaining book for the general public, has something to do with the rather sweeping--and sometimes sloppy--arguments for Bayesian over frequentist methods. Although Silver gives a few engaging and precise examples of Bayes' rule in operation (playing poker or deciding whether your domestic partner is cheating on you, for instance), he is quick to characterize a variety of informal, intuitive modes of combining many different kinds of data as tantamount to following Bayes' rule. Marcus & Davis (2013) identify one telling example--a very successful sports bettor who recognizes the importance of data that the bookies overlook, misjudge, or do not acquire . (Pp. 232-61). What makes this gambler a Bayesian? Silver thinks it is the fact that "[s]uccessful gamblers ... think of the future as speckles of probability, flickering upward and downward like a stock market ticker to every new jolt of information." (P. 237). That's fine, but why presume that the flickers follow Bayes' rule as opposed to some other procedure for updating beliefs? And why castigate frequentist statisticians, as Silver seems to, as "think[ing] of the future in terms of no-lose bets, unimpeachable theories, and infinitely precise measurements"? Ibid. Surely, that is not the world in which statisticians live.

Changing probability judgments does not make someone a Bayesian

In proselytizing for Bayes' theorem and in urging readers to "think probabilistically" (p. 448), Silver also writes that
When you first start to make these probability estimates, they may be quite poor. But there are two pieces of favorable news. First, these estimates are just a starting point: Bayes's theorem will have you revise and improve them as you encounter new information. Second, there is evidence that this is something we can learn to improve. The military, for instance, has sometimes trained soldiers in these techniques,5 with reasonably good results.6 There is also evidence that doctors think about medical diagnoses in a Bayesian manner.7 [¶] It is probably better to follow the lead of our doctors and our soldiers than our television pundits.
It is hard to argue with the concluding sentence, but where is the evidence that many soldiers and doctors are intuitive (or trained) Bayesians? The report cited (n.5) for the proposition that "[t]he military ... has sometimes trained soldiers in the [Bayesian] techniques" says nothing of the kind.* Similarly, the article that is supposed to show that the alleged training in Bayes' rule produces "reasonably good results" is quite wide of the mark. It is a 35-year-old report for the Army on "Training for Calibration" about research that made no effort to train soldiers to use Bayes' rule.**

How about doctors? The source here is an article in the British Medical Journal that asserts that "[c]linicians apply bayesian reasoning in framing and revising differential diagnoses." Gill et al. (2005). But these authors--I won't call them researchers because they did no real research--rely only on their impressions and post hoc explanations for diagnoses that are not expressed probabilistically. As one distinguished physician tartly observed, "[c]linicians certainly do change their minds about the probability of a diagnosis being true as new evidence emerges to improve the odds of being correct, but the similarity to the formal Bayesian procedure is more apparent than real and it is not very likely, in fact, that most clinicians would consider themselves bayesians." Waldron (2008, pp. 2-3 n.2).

The transposition fallacy

Consistent with this tendency to conflate expressing judgments probabilistically with using Bayes' rule to arrive at the assessments, Silver presents probabilities that have nothing to do with Bayes' rule as if they are properly computed posterior probabilities. In particular, he naively transposes conditional probabilities to misrepresent p-values as degrees of belief.

At page 185, he writes that
A once-famous “leading indicator” of economic performance, for instance, was the winner of the Super Bowl. From Super Bowl I in 1967 through Super Bowl XXXI in 1997, the stock market gained an average of 14 percent for the rest of the year when a team from the National Football League (NFL) won the game. But it fell by almost 10% when a team from the original American Football Leage (AFL) won instead. [¶] Through 1997, this indicator had correctly “predicted” the direction of the stock market in twenty-eight of thirty-one years. A standard test of statistical significance, if taken literally, would have implied that there was only about a 1-in-4,700,000 possibility that the relationship had emerged from chance alone.
This is a cute example of the mistake of interpreting a p-value, acquired after a search for significance, as if there had been no such search. As Silver submits, "[c]onsider how creative you might be when you have a stack of economic variables as thick as a phone book." Ibid.

But is the ridiculously small p-value (that he obtained by regressing the S&P 500 index on the conference affiliation of the Super Bowl winner) really the probability "that the relationship had emerged from chance alone"? No, it is the probability that such a remarkable association would be seen if the Super Bowl outcome and the S&P 500 index were entirely uncorrelated (and no one had searched for a data set that shared a seemingly shocking correlation to the S&P 500 index). Silver may be a Bayesian at heart, but he did not compute the probability of the null hypothesis given the data, and it is problematic to tell the reader that a "standard test of statistical significance" (or more precisely, a p-value) gives the (Bayesian) probability that the null hypothesis is true.

Of course, with such an extreme p-value, the misinterpretation might not make any practical difference, but the same misinterpretation is evident in Silver's characterization of a significance test at the 0.05 level. He writes that "[b]ecause 95 percent confidence in a statistical test is Fisher’s traditional dividing line between 'significant' and 'insignificant,' researchers are much more likely to report findings that statistical tests classify as 95.1 percent certain than those they classify as 94.9 percent certain—a practice that seems more superstitious than scientific." (P. 256 n. †). Sure, any rigid dividing line (a procedure that Fisher did not really use) is arbitrary, but rejecting a hypothesis in a classical statistical test at the 0.05 level does not imply a 95% certainty that this rejection is correct.

In transposing conditional probabilities in violation of both Bayesian and frequentist precepts, Silver is in good and plentiful company. As the pages on this blog reveal, theoretical physicists, epidemiologists, judges, lawyers, forensic scientists, journalists, and many other people make this mistake. E.g., The Probability that the Higgs Boson Has Been Discovered, July 6, 2012. Despite its general excellence in describing data-driven thinking, The Signal and the Noise would have benefited from a little more error-correcting code.

Notes

* Rather, Gunzelmann & Gluck (2004) discusses training in unspecified "mission-relevant skills." An "expert model is able to compare their actions against the optimal actions in the task situation" and "identify [trainee errors] and provide specific feedback about why the action was incorrect, what the correct action was, and what the students should do to correct their mistake." The expert model--and not the soldiers--"uses Bayes’ theorem to assess mastery learning based upon the history of success and failure on particular units of skill within the task." Ibid. According to the authors, this "Bayesian knowledge tracing approach" is inadequate because it "does not account for forgetting, and thus cannot provide predictions about skill retention." Ibid.

** Lichtenstein & Fischhoff's (1978) objective was "to help analysts to more accurately use numerical probabilities to indicate their degree of confidence in their decisions." They did not study military analysts, but instead recruited 12 individuals from their personal contacts. They had these trainees assess the probabilities of statements in the areas of geography, history, literature, science, and music. They measured how well calibrated their subjects were. (A well calibrated individual gives correct answers to x% of the questions for which he or she assesses the probability of the given answer to be x%.) The proportion of the subjects whose calibration improved after feedback was 72%. Ibid.

References

Walter Frick, Nate Silver on Finding a Mentor, Teaching Yourself Statistics, and Not Settling in Your Career, Harvard Business Review Blog Network, Sept. 24, 2013, http://blogs.hbr.org/2013/09/nate-silver-on-finding-a-mentor-teaching-yourself-statistics-and-not-settling-in-your-career/.

Christopher J. Gill, Lora Sabin & Christopher H. Schmid, Why Clinicians Are Natural Bayesians, 330 Brit. Med. J. 1080–83 (2005), available at http://www.ncbi.nlm.nih.gov/pmc/articles/PMC557240/

Glenn F. Gunzelmann & Kevin A. Gluck, Knowledge Tracing for Complex Training Applications: Beyond Bayesian Mastery Estimates, in Proceedings of the Thirteenth Conference on Behavior Representation in Modeling and Simulation 383-84 (2004), available at http://act-r.psy.cmu.edu/wordpress/wp-content/uploads/2012/12/710gunzelmann_gluck-2004.pdf.

Sarah Lichtenstein & Baruch Fischhoff, Training for Calibration, Army Research Institute Technical Report TR-78-A32, Nov. 1978, available at http://www.dtic.mil/dtic/tr/fulltext/u2/a069703.pdf

Gary Marcus & Ernest Davis, What Nate Silver Gets Wrong, New Yorker, Jan. 25, 2013, http://www.newyorker.com/online/blogs/books/2013/01/what-nate-silver-gets-wrong.html

Tony Waldron, Palaeopathology (2008), excerpt available at http://assets.cambridge.org/97805216/78551/excerpt/9780521678551_excerpt.pdf

Thursday, 16 January 2014

Tres Mal Errors with DNA Evidence

A story in the Denver Post (Gurman 2014) begins with the disturbing news that
A malfunction in a DNA processing machine led to the scrambling of samples from 11 Denver police burglary cases, officials acknowledged Friday. It took more than two years for the department to discover the errors. As a result of the mix-up, prosecutors are dismissing burglary cases against four people, three of whom had already pleaded guilty.
What happened?

In 2011, "[a] machine 'froze' while running a tray of 19 DNA samples." An analyst "replaced [the samples] in the wrong order" after asking the manufacturer of the robot how to proceed. More than two years later, "the machine froze for a second time." An analyst called again and "became concerned because the directions seemed different the second time. Further review over the next month revealed" the 2011 error. Ibid.

What of It?

The police department chief of staff observed that "[n]one of the DNA was compromised; it was merely associated with the wrong case when we were done." Ibid. In other words, the crime-scene DNA profiles were mislabeled. Such errors could have helped criminals avoid detection. For instance, a burglar in case A falsely associated with case B might have had a strong alibi defense for case B. Alternatively, such labeling errors could have caused individuals to be convicted of the wrong crime -- perhaps a man guilty of a burglary could have been found guilty of an murder (or vice versa).

In this incident, however, a police spokeswoman said that "[a]ll four people had confessed to at least one burglary, but the DNA error meant they were charged with the wrong ones." Ibid. Prosecutors dismissed the charges against the four, and they will not be tried for the other burglaries because the statute of limitations has expired.

Another "tray mal" case

Misuse of automated machinery for DNA analysis also produced an error--this one involving an entirely innocent man--in "what is described as the most advanced automated DNA testing system in the UK at LGC forensics labs in Teddington." Israel 2012. The machinery extracts DNA from wells in a plastic tray. Police arrested Andrew Scott, 20, after a street fight and sent a saliva to LGC for profiling. (Doyle 2012). Instead of throwing away the tray after the run with Scott's saliva sample, however, a worker reused it in an unrelated rape case. The tray contained enough of Scott's left-over DNA to show his DNA profile in the later rape sample. The laboratory should have been aware of a problem, for "[t]he batch containing the rape sample showed DNA present in the negative control (a blank sample put through to test for contamination)." (Rennison 2012).

As a result of the error, Scott was charged with "a violent attack on a woman in Manchester – carried out when he was hundreds of miles away in Plymouth." (Doyle 2012). After he spent months in prison, the charge was dismissed. "Phone records showed he was 300 miles away on the south coast when the rape took place." Scott described the experience as a "living nightmare": "They kept me in a segregation wing which was full of rapists and paedophiles. I suffered lots of verbal abuse and other inmates spitting at us and shouting 'paedos.'" Ibid.

References
Acknowledgments
  •  Thanks to Bill Thompson for alerting me to the Denver case.
Copyr. (c) DH Kaye 2014

Tuesday, 24 December 2013

Breathalyzers and Beyond: The Unintuitive Meanings of "Measurement Error" and "True Values" in the 2009 NRC Report on Forensic Science

Five years ago, the National Research Council released its eagerly awaited and repeatedly postponed report on "Strengthening Forensic Science in the United States: A Path Forward." One theme of the report was that forensic experts must present their findings with due recognition of Rumsfeldian "known unknowns." For example, the report repeatedly referred to "the importance of ... a measurement with an interval that has a high probability of containing the true value" (NRC Committee 2009, p. 121), and it referred to "error rates" for categorical determinations (ibid., pp. 117-22). 

Earlier this year, UC-Davis law professor and evidence guru Edward Imwinkelried and I submitted a letter urging the Washington Supreme Court to review a case raising the issue of whether the state courts should admit point estimates of blood or breath alcohol concentration without an accompanying quantitative estimate of the uncertainty in each estimate. (The court denied review.) Since the NRC report uses breath-alcohol measurements to explain the meaning of its call for interval estimates, one would think that the report would have a good illustration of a suitable interval. But that is not what I found. The report's illustration reads as follows:
As with all other scientific investigations, laboratory analyses conducted by forensic scientists are subject to measurement error. Such error reflects the intrinsic strengths and limitations of the particular scientific technique. For example, methods for measuring the level of blood alcohol in an individual or methods for measuring the heroin content of a sample can do so only within a confidence interval of possible values. In addition to the inherent limitations of the measurement technique, a range of other factors may also be present and can affect the accuracy of laboratory analyses. Such factors may include deficiencies in the reference materials used in the analysis, equipment errors, environmental conditions that lie outside the range within which the method was validated, sample mix-ups and contamination, transcriptional errors, and more.

Consider, for example, a case in which an instrument (e.g., a breathalyzer such as Intoxilyzer) is used to measure the blood-alcohol level of an individual three times, and the three measurements are 0.08 percent, 0.09 percent, and 0.10 percent. The variability in the three measurements may arise from the internal components of the instrument, the different times and ways in which the measurements were taken, or a variety of other factors. These measured results need to be reported, along with a confidence interval that has a high probability of containing the true blood-alcohol level (e.g., the mean plus or minus two standard deviations). For this illustration, the average is 0.09 percent and the standard deviation is 0.01 percent; therefore, a two-standard-deviation confidence interval (0.07 percent, 0.11 percent) has a high probability of containing the person’s true blood-alcohol level. (Statistical models dictate the methods for generating such intervals in other circumstances so that they have a high probability of containing the true result.)
(Ibid., pp. 116-17.)

What is troublesome about this explanation? Let me count the ways.

1. "Measurement error" does not refer to all errors of measurement

"[D]eficiencies in the reference materials used in the analysis, equipment errors, environmental conditions that lie outside the range within which the method was validated, sample mix-ups and contamination, transcriptional errors, and more" all "can affect the accuracy of laboratory analyses." Nevertheless, they do no count as "measurement error" because they are "factors other than the inherent limitations of the measurement technique." Not being "intrinsic [to] the particular scientific technique," they fall outside the committee's definition of "measurement error."

That narrow definition calls to mind the claims of some fingerprint analysts that the ACE-V method has an "methodological" error rate of zero because the only possibility for error arises when a human being does not apply the method perfectly. The difference, however, is that one can measure the errors when the breathalyzer has no deficient reference materials, no extreme environmental conditions, no sample mix-ups and contamination, no transcriptional errors, and so on. The fingerprint analyst, in contrast, is the measuring instrument, and it is impossible to distinguish between instrument measurement error and human error in that context.

There is nothing illogical in quantifying some but not all measurement errors when some are more readily and validly quantifiable than others. Machines might not be tested periodically to ensure that they are operating as they are supposed to (e.g., DiFilipo 2011; Sovern 2012), but whether one can usefully build that possibility into the computation of the uncertainty of a measurement that might be suitable for courtroom testimony is not clear. Yet, using the seemingly all-encompassing phrase "measurement error" in a narrow, technical sense -- to denote only the noise inherent in the apparatus when operated under certain conditions -- is potentially misleading.

2. "True values" are not true blood-alcohol levels.

Because the committee's example of "measurement error" quantifies only "intrinsic" error, its statement that "a two-standard-deviation confidence interval (0.07 percent, 0.11 percent) has a high probability of containing the person’s true blood-alcohol level" also is easily misunderstood. The confidence interval (CI) for "true values" does not pertain to the actual blood-alcohol level. That level can differ from the point estimate of 0.09 for other reasons, making the real uncertainty greater than ± 0.02.

In addition, a breathalyzer measures alcohol in the breath, not in the bloodstream. The concentrations are related, but the precise functional relationship varies across individuals (e.g., Martinez & Martinez 2002). This is another source of uncertainty not reflected in the committee's CI for blood-alcohol concentration (BAC), although the committee could have sidestepped this issue by referring to breath-alcohol concentration (BrAC).

3. The standard error of the breathalyzer would be determined differently.

The NRC committee imagines using a breathalyzer to make three measurements of the same breath sample. The parenthetical, concluding sentence about "statistical models" for "other circumstances" suggests that the committee realized that this approach is not one that anyone would use to estimate the noise in the apparatus. The breathalyzer should be tested on many samples with known concentrations to ensure that it is not biased and to quantify the extent of the random variations about those known values. Manufacturers perform such tests (e.g., Coyle et al. 2010).

4. A CI of ±2 standard errors might not have "a high probability of containing the person’s true blood-alcohol level"

Let's put aside all the concerns raised so far. Suppose that the errors in the machine's measurement always are normally distributed about the true value in a breath sample; that the applicable standard deviation for this distribution is 0.01; and that the single measured value is 0.09. Is it now true that the interval 0.09 ± 0.02 "has a high probability of containing the person’s true blood-alcohol level"?

Maybe. Two standard errors give an interval with a confidence coefficient of approximately 95%. That is to say that this one interval comes from a procedure that generates intervals that cover the true value about 95% of the time. It is tempting to say that the probability that the interval in question covers the true value therefore is 95%.

But let's think about how the sample came to be tested. The arrested officer picks someone out of a population of motorists. The motorists have varying levels of BrACs, and the officer has some level of skill in spotting the ones who might well be inebriated. Suppose that the drivers the officer stops and tests have BrACs that are normally distributed with mean 0.04 and standard deviation 0.01. The officer's breathalyzer is functioning according to manufacturer's specifications, and the standard deviation in its measurements is 0.01, as in the NRC report. Having obtained a measurement of 0.08 on the one driver's breath sample, what is a high probability interval for true BrAC in this one breath sample? Is it 0.07 to 0.11?

It turns out the probability that the true BrAC falls within the NRC's interval is only 24% (applying equations 2.9 and 2.10 in Gelman et al. 2004). If the officer stopped drivers who whose mean BrAC were greater than 0.04 or with more variable BrACs, the probability for the NRC's interval being correct would be greater. If, for example, the standard deviation in this group were 0.02 instead of 0.01 (and the mean were still 0.04), then the probability for the NRC's interval would be 87%.

Of course, we do not know much about the distribution of BrAC in the group that the officer stops. As indicated above, this distribution would depend on the drinking habits of drivers in the town and the officer's skill in pulling over drunken drivers. The choice of a normal distribution with the parameters mentioned above is not likely to be realistic. But whatever the distribution may be, it, along with the single measured value, bears on the true value of the tested driver's BrAC. This fact makes it tricky to quantify the probability that the NRC's CI includes the driver's BrAC.

* * *

The NRC Report was certainly correct to call on forensic scientists to develop better measures of the uncertainty in their findings and to apply them in their reports and testimony. But figuring out what these measures should be and how to use them is a formidable challenge. Meeting this challenge will be a lot harder than the simple example of a confidence interval in the report might suggest.

References

Sunday, 8 December 2013

Error on Error: The Washington 23

Frequently cited in warnings on the risks of errors in DNA typing is a 2004 article prepared by unnamed staff of the Seattle Post-Intelligencer. In one highly praised book, for instance, Sheldon Krimsky of Tufts University and Tania Simoncelli, then with the ACLU, wrote that the paper “reported that forensic scientists at the Washington State Patrol Laboratory had made mistakes while handling evidence in at least 23 major criminal cases over three years” [1, p. 280]. The article itself begins “[c]ontamination and other errors in DNA analysis have occurred at the Washington State Patrol crime labs, most of it the result of sloppy work” [2].

Laboratory documentation of “sloppy work” should be encouraged. It should be scrutinized inside and outside of the laboratory. Within the laboratory, it can be a path to improvements. Outside the laboratory world, reporting on problems, quotidian and catastrophic alike, can increase the level of public and professional understanding of how forensic science is practiced. However, it is important to be clear about the nature, severity, and implications of specific “mistakes,” “errors,” and “contamination.” These terms cover a variety of phenomena.

Even before the earliest days of PCR-based DNA typing, it has been known that “contamination” is an omnipresent possibility. It can result from extraneous DNA in materials from companies that supply reagents and equipment, from the introduction of the analyst’s DNA into the sample being analyzed (“for example, when the analyst talks while handling a sample, leaving an invisible deposit of saliva” [2]), from inadequate precautions against transferring DNA from one test with one sample over to another test with a different sample (a form of “cross-contamination”), and so on. Many forms of contamination are detectable, but they can complicate or interfere with the interpretation of an STR profile [3]. Cross-contamination of a crime-scene sample with a potential suspect’s DNA either before or after it reaches the laboratory is particularly serious because it could result in a false match.

As described in an appendix below, it appears that only one of the 23 cases (#22) involved a false report of a match, and the report was corrected before any charges were filed. However, Bill Thompson presented a different case as a premier example of "false cold hits" [4, p. 230]. In his latest publication on errors in DNA typing, he wrote that
[W]hile the Washington State Crime Patrol Laboratory a cold-case investigation of a long-unsolved rape, it found a DNA match to a reference sample in an offender database, but it was a sample from a juvenile offender who would have been a toddler at the time the rape occurred. This prompted an internal investigation at the laboratory that concluded that DNA from the offender's sample, which had been used in the laboratory for training purposes, had accidentally contaminated samples from the rape case, producing a false match. [4, p. 230].
Thompson noted that he "assisted the newspaper in the investigation" [4, p. 341 n.12]. Apparently, he was referring to case #5 in the article (although the article labels it a homicide case). In any event, it is the only case Thompson lists as an example of a false match in Washington.

My conclusion is that the Washington cases certainly establish that mistakes of many types can occur in DNA laboratories and that some types of mistakes can produce false matches, false accusations, and even false convictions. But none of the 23 are themselves instances of false charges or false convictions. This conclusion neither condones the mistakes nor excludes the possibility that DNA has produced such outcomes in Washington.  But it may help put the 23 cases and the writing about them in perspective.

References
  1. Sheldon Krimsky & Tania Simoncelli, Genetic Justice: DNA Data Banks, Criminal Investigations, and Civil Liberties (2011)
  2. DNA Testing Mistakes at the State Patrol Crime Labs, Seattle Post-Intelligencer, July 21, 2004, 10:00 pm, http://www.seattlepi.com/local/article/DNA-testing-mistakes-at-the-State-Patrol-crime-1149846.php
  3. Terri Sundquist & Joseph Bessetti, Identifying and Preventing DNA Contamination in a DNA-Typing Laboratory, Profiles in DNA, Sept. 2005, at 11-13, http://www.promega.com/~/media/Files/Resources/Profiles%20In%20DNA/802/Identifying%20and%20Preventing%20DNA%20Contamination%20in%20a%20DNA%20Typing%20Laboratory.ashx
  4. William C. Thompson, The Myth of Infallibility, in Genetic Explanantions: Sense and Nonsense 227 (Sheldon Krimsky & Jeremy Gruber eds. 2013)
Related postings
APPENDIX
23 and Me

This Appendix quotes the newspaper descriptions in full, then offers my own remarks.

EXAMPLE NO. 1
Problem: Cross-contamination
When and where: July 2002, Spokane lab
Forensic scientist: Lisa Turpen
Case: child rape
What happened: Turpen contaminated one of four vaginal swabs with semen from a positive control sample. Corrected report issued almost two years later in March 2004. ....Yakima prosecutors offered plea deal during the trial, with defendant pleading guilty to two gross misdemeanors. Turpen's mistake was a factor, according to defense.”

REMARKS: I do not know what “semen from a positive control sample” means. When DNA from a cell line is used to ensure that PCR is amplifying those alleles, the cell-line DNA is known as a positive control sample. This example does not sound like a case of contamination involving that kind of a positive control. Adding semen to a vaginal swab obviously is unacceptable, but if the other three swabs produced a single male DNA profile and the fourth showed two male profiles in a case involving a single rapist, the anomalous profile would not be falsely matched to anyone.

EXAMPLE NO. 2
Problem: Erroneous lab report
When and where: August 2002, Seattle lab
Forensic scientist: William Stubbs
Case: Fatal police shooting of Robert Thomas
What happened: Two hours before testifying at inquest, Stubbs discovered his crime lab report was wrong and notified prosecutor. His report said test found brown stain on gun was likely blood, but his notes had no indication of blood. ... Corrected report issued in September 2002. ... Co-worker reviewing case did not catch mistake.

REMARK: Does not involve DNA typing.

EXAMPLE NO. 3
Problem: Self-contamination
When and where: April 2001, Spokane lab
Forensic scientists: Charles Solomon, Lisa Turpen
Case: rape/kidnapping/assault
What happened: In separate tests, Solomon and Turpen contaminated hair-root tests with their own DNA. Solomon also contaminated reference blood sample with his DNA. ...Three defendants were convicted.

REMARK: There is no suggestion of a false match here.

EXAMPLE NO. 4
Problem: Testing error
When and where: September 2002, Marysville lab
Forensic scientist: Mike Croteau
Case: robbery/assault
What happened: Rushing to meet deadlines, Croteau mixed up reference samples from victim and suspect. He reported incorrect findings verbally to prosecutor, then discovered his mistake. ... Defendant pleaded guilty.

REMARK: What is the mistake here? It must be something more than using the wrong names for the two samples that were compared to produce a false match.

EXAMPLE NO. 5
Problem: Cross-contamination
When and where: August 2003, Seattle lab
Forensic scientist: Robin Bussoletti
Case: homicide
What happened: Bussoletti likely contaminated work surface while testing a blood sample from a convicted felon during training. Next DNA analyst who used work station noticed contamination in chemical solution that is not supposed to contain DNA.

REMARKS: Definitely sloppy -- and potentially falsely incriminating if work surface was then used without a thorough cleaning for casework.

EXAMPLE NO. 6
Problem: Cross-contamination
When and where: January 2004, Tacoma lab
Forensic scientist: Jeremy Sanderson
Case: child rape
What happened: Sanderson failed to change gloves between handling evidence in two cases. He noticed contamination in chemical solution. ... Defendant convicted and sent to prison.

REMARK: Is this a case of cross-contamination of samples?

EXAMPLE NO. 7
Problem: Error during testing
When and where: June 2002, Seattle lab
Forensic scientist: Denise Olson
Case: aggravated murder
What happened: Olson did initial test to look for blood on shoes. She got weak positive result, then threw out swabs. She didn't document findings or notify police. Kirkland police complained because discarded swabs couldn't be tested for DNA. ... Shoes sent to private lab for retesting. ... Defendant Kim Mason convicted and sentenced to life without release.

REMARK: Not a false match

EXAMPLE NO. 8
Problem: Error in DNA test interpretation
When and where: October 1998, Seattle lab
Forensic scientist: George Chan
Case: rape
What happened: Chan misstated statistical likelihood of match with suspect. Co-worker reviewing case didn't catch error. ... Pierce County prosecutor noticed mistake at pretrial conference in September 2000. ... Defendant convicted.

REMARK: Not a false match

EXAMPLE NO. 9
Problem: Error in testing procedure
When and where: September 2002, Seattle lab
Forensic scientist: Denise Olson
Case: robbery/assault
What happened: Olson tested known DNA samples before evidence collected at crime scene -- a violation of lab procedure aimed at preventing cross-contamination. A co-worker caught the mistake while reviewing the case.... Tests were redone. ... Defendant pleaded guilty.

REMARK: This departure from protocol raises the risk of an incriminating case of cross-contamination, but there is no indication that any cross-contamination occurred.

EXAMPLE NO. 10
Problem: Self-contamination
When and where: November 2002, Tacoma lab
Forensic scientist: Mike Dornan
Case: rape

What happened: Dornan contaminated DNA test of victim's underwear with his own DNA. May have resulted from talking during testing process.... Defendant pleaded guilty.

REMARK: No false match.

EXAMPLE NO. 11
Problem: Unknown source of contamination
When and where: January 2004, Tacoma lab
Forensic scientist: Christopher Sewell
Case: homicide
What happened: Sewell found low level of DNA from unknown source in blood sample from victim. May have come from blood transfusion of victim before death. ... Case pending.

REMARK: The “unknown source of contamination” does not seem to have produced a false match if peak heights indicated a minor contributor, and the major contributor was the defendant,

EXAMPLE NO. 12
Problem: Self-contamination
When and where: March 2004, Tacoma lab
Forensic scientist: William Dean
Case: rape
What happened: Dean contaminated control sample with his own DNA while testing police evidence. ... No suspect.

REMARK: No suspect, no contamination of a crime-scene or suspect sample, no false match.

EXAMPLE NO. 13
Problem: Unknown source of contamination
When and where: January 2003, Spokane lab
Forensic scientist: Lisa Turpen
Case: murder
What happened: Turpen found unidentified female DNA in control sample while testing evidence in Stevens County double-murder case.... Defendant convicted.

REMARK: No contamination of a crime-scene or suspect sample, no false match.

EXAMPLE NO. 14
Problem: Unknown source of contamination
When and where: January 2003, Spokane lab
Forensic scientist: Lisa Turpen
Case: robbery/kidnapping
What happened: Turpen found unidentified female DNA in control sample while testing evidence in Yakima County case. Evidence tested same day as evidence in Example No.13.... Case pending.

REMARK: No contamination of a crime-scene or suspect sample, no false match.

EXAMPLE NO. 15
Problem: Self-contamination
When and where: September 2003, Marysville lab
Forensic scientist: Greg Frank
Case: murder
What happened: Frank contaminated control samples with his own DNA during testing in Snohomish County case. ...Case pending.

REMARK: No contamination of a crime-scene or suspect sample, no false match.

EXAMPLE NO. 16
Problem: Self-contamination
When and where: September 2003, Marysville lab
Forensic scientist: Greg Frank
Case: child molestation/rape
What happened: Frank contaminated control samples with his own DNA during testing in Kitsap County case. ... Defendant pleaded guilty.

REMARK: No contamination of a crime-scene or suspect sample, no false match.

EXAMPLES NO. 17 & 18
Problem: Unknown source of contamination
When and where: October 2003, Seattle lab
Forensic scientists: Phil Hodge, Amy Jagman
Cases: unknown
What happened: Hodge and Jagman both discovered unknown source of contamination in chemical used during DNA testing. Chemical discarded and evidence retested.

REMARK: No contamination of a crime-scene or suspect sample, no false match.

EXAMPLE NO. 19
Problem: Self-contamination
When and where: October 2002, Spokane lab
Forensic scientists: Charles Solomon, Lisa Turpen
Case: murder
What happened: Solomon found Turpen's DNA on three bullet casings retrieved from scene of Richland double murder. ... Defense expert disputed this at trial, testifying that DNA profile belonged to unknown female. ... Defendant Keith Hilton convicted.

REMARK: No false match.

EXAMPLE NO. 20
Problem: Cross-contamination
When and where: February 2002, Tacoma
Forensic scientist: Mike Dornan
Case: child rape
What happened: Dornan contaminated evidence in King County rape case with DNA from a previous case, likely by failing to properly sterilize scissors. ... Defendant pleaded guilty to a reduced charge before contamination was discovered.

REMARK: I presume that if the previous case were the defendant’s and that is what led to the charge against the defendant, the newspaper would have so stated. That would have been a false match.

EXAMPLE NO. 21
Problem: Self-contamination
When and where: January 2001, Marysville lab
Forensic scientist: Brian Smelser
Case: rape
What happened: Smelser contaminated three tests with his own DNA in Kirkland rape case. Prosecutor had to send remaining half-sample to California lab for retesting.... Defendant pleaded guilty to reduced charge.

REMARK: No false match.

EXAMPLE NO. 22
Problem: Error in testing
When and where: December 2002, Seattle lab
Forensic scientist: Denise Olson
Case: rape/attempted murder
What happened: Olson misinterpreted DNA results, telling Seattle police their suspect was a match. Co-worker caught error 11 days later, just as charges were about to be filed.... Case unsolved.

REMARK: A false positive report (not resulting from contamination).

EXAMPLE NO. 23
Problem: Self-contamination
When and where: January 2004, Seattle lab
Forensic scientist: George Chan/William Stubbs
Case: child rape
What happened: Chan's DNA found in suspect's boxer shorts by Stubbs. Problem traced to Chan talking to Stubbs during testing.... Suspect pleaded guilty.

REMARK: No contamination of a crime-scene or suspect sample, no false match.

Saturday, 7 December 2013

Error on Error: Quashing Brian Kelly's Conviction

Are there any errors in DNA testing? Are there any errors that produce false positives? Do DNA databases generate any false leads? Do false leads produce any false arrests? Any false convictions? The answers to these questions are yes, yes, yes, yes, and yes. (See related postings below.)

But how large is the risk of a false positive match to an existing suspect? To an innocent individual culled from a database? By and large, we are limited to isolated reports in newspapers--reports that are newsworthy precisely because they are rare. The most complete compilation of the troubling cases, presented in a survey of the ways that errors can arise, is to be found in a book chapter by Bill Thompson of the University of California at Irvine. [1]

Professor Thompson is an unusually knowledgeable and astute commentator, consultant, and advocate in the field of DNA evidence, and it should be revealing to work through his examples. That is what I have started to do. So far, I have looked into only the very first case noted in the chapter. According to Professor Thompson, it exemplifies a "common problem" [1, p. 230] of "[a]ccidental transfer of cellular material or DNA from one sample to another" [1, p. 229] causing "false reports of a DNA match between samples that originated from different people" [1, p. 230].

The example is a 1988 DNA test in a rape case in Scotland that led to the conviction of Brian Kelly. Thompson simply reports that "Scotland's High Court of Justiciary quashed a conviction in one case in which the convicted man (with the help of sympathetic volunteer scientists) presented persuasive evidence that the DNA match that incriminated him arose from a laboratory accident" [1, 230]. The "accident" in question consisted of DNA leaking from one well to an adjacent one in an agarose gel used in VNTR typing or an analyst's misloading some of the same DNA sample into both wells instead of just the one she was aiming for.

But the evidence that Professor Thompson found "persuasive" did not persuade the court. Indeed, the experts did not even testify that leakage or misloading had occurred. Rather, they stated that it was a "low risk" event, that the possibility could not be excluded, and that a procedure that would have reduced the risk could have been followed (and was adopted two years later) [2, ¶¶ 15-17].

Thus, the Scottish Appeals Court, noting other evidence in Kelly's favor and weaknesses in the Crown's case, quashed the conviction--but not because it concluded that that the match was false. The court quashed the conviction because the jury was not informed of the fact that the same DNA could end up in two adjacent lanes. The court wrote:
It was not suggested that there is evidence positively indicating that cross-contamination did, or may have, occurred. On the basis of the evidence tendered by the appellant, it is maintained, on the other hand, that there was a risk of cross-contamination arising from the practice at that time of using adjoining wells for DNA samples from the crime scene and the suspect, and of such cross-contamination being undetected. It was not in controversy that it was possible for there to be leakage between adjoining wells or for DNA material to fall accidentally into a well next to the one for which it was intended. Up to a point the evidence ... as to the procedures which were followed, and the special care which was taken, countered the risk that such a mishap would in practice occur or be undetected. However, such evidence does not in our view provide a complete answer. In particular there was, on the evidence, a risk that the leakage of DNA from the well for the suspect's reference sample to the adjoining well which already held the crime scene sample would not be detected. It was, of course, a low risk, but it was of sufficient importance to be recognised by experts ... .

... In our opinion there is evidence which is capable of being regarded as credible and reliable as to the existence of a risk of cross-contamination occurring without it being detected. The risk was a low risk. It may be that in other circumstances the fact that the jury did not hear such evidence would not lead to the conclusion that there had been a miscarriage of justice. However, in the present case it is otherwise since the DNA evidence was plainly of critical importance for the conviction of the appellant. If the jury had rejected that evidence there would, in our view, have been insufficient evidence to convict the appellant. Accordingly, while the evidence related to a low risk of cross-contamination, the magnitude of the implications for the case against the appellant were substantial. For these reasons we have come to the conclusion that the appellant has established the existence of evidence which is of such significance that the fact that it was not heard by the jury constituted a miscarriage of justice. [2, ¶ 21-22]

Based on this opinion, the 1988 DNA testing with a superseded technology is a far cry from is a true example of an innocent man convicted because of "a laboratory accident." It is nothing more--or less--than a case in which the defendant did not present expert testimony at trial that the laboratory used a procedure that left open a preventable mode of cross-contamination. The case is an appropriate illustration of the importance of improving laboratory practices, but such cases are not proof of known "false reports" commonly resulting from cross-contamination.

References
  1. William C. Thompson, The Myth of Infallibility, in Genetic Explanantions: Sense and Nonsense 227 (Sheldon Krimsky & Jeremy Gruber eds. 2013)
  2. Opinion in the Reference by the Scottish Criminal Cases Review Commission in the Case of Brian Kelly, Appeal Court, High Court of Justiciary, Appeal No. XC458/03, Aug. 6, 2004, http://www.scotcourts.gov.uk/opinions/XC458.html
Related postings