Thursday, 12 June 2014

Flawed Journalism on Flawed Forensics in Slate Magazine

Yesterday, Slate magazine published an article by Mark Joseph Stern announcing that “Forensic Science Isn’t Science.” 1/ The writer’s objective—to urge that forensic science be conducted rigorously and fairly—is laudable. But just as shabby science should not be tolerated in the courtroom or the police station, journalism that pays little heed to the facts should not be acceptable in serious publications.

I’ll give one example, chosen because it pervades the publication. The article begins with the claim that “[f]orensic analysis of semen introduced at trial had convinced the jury that [Earl] Washington [Jr.] ... had brutally raped and murdered a young woman in 1982.” It asks, “[h]ow could forensic evidence, widely seen as factual and unbiased, nearly send [this] innocent person to his death?” It ends with the plaintive thought that “[o]ur national experiment in untested forensics may soon be coming to a close. But it hasn’t ended in time to prevent a few more people like Earl Washington from being sacrificed on the altar of pseudoscience.”

The conviction and exoneration of Earl Washington have much to teach us about criminal justice. But it would be hard to find a worse example of “an innocent man being sacrificed on the altar of pseudoscience.”  There was no forensic evidence—scientific or pseudoscientific—introduced in the trial. Had there been, the outcome might have been different. This is the conclusion that follows from the description of the case in an important book, Convicting the Innocent, by Professor Brandon Garrett.

Garrett's research reveals that the police made every effort to keep science away when they built their case around a classic false confession from a “borderline mentally retarded farmhand” 2/ with convincing detail fed to him by police. One officer was found in a later civil rights action to have “fabricated the confession.” 3/ The alleged confession included the revelation (known to the police) of the killer’s blood-stained shirt with a torn-off patch left in the victim’s dresser drawer. Although forensic analysts had excluded five other suspects as possible sources of hairs found in the shirt pocket, “police instructed the state crime laboratory not to compare [Washington’s] hairs.” 4/

Even more telling—but untold—was the serological evidence in the case. According to Mr. Stern, it was “semen introduced at trial” that “convinced the jury.” But no semen was introduced at trial. No “semen analysis,” as Mr. Stern calls it, was offered into evidence. If only it had been!

“The semen-stained blanket from the victim’s bed was blood-typed, and that rudimentary technique had ruled out Washington.” 5/ The prosecutor would hardly want to introduce this evidence. (Indeed, it is hard to see how he ethically could go to trial without having proof that the blood-typing was incorrect.) As for the inexperienced defense counsel, 6/ “[t]he lawyer later said that while he saw the forensic reports, he ‘was not familiar with the significance of the analysis.’” 7/ Worse still, “the state concealed crucial evidence of innocence, including forensic evidence, from the defense.” 8/

In short, presenting the conviction and near-execution of Earl Washington, Jr., as the example of “a decades-long experiment in which undertrained lab workers jettison the scientific method in favor of speedy results that fit prosecutors’ hunches” disguises the real lessons of the case. The Washington case is an awful illustration of (1) evidence of a false confession that could have been prevented by proper interviewing techniques (including recording the confession); (2) willful blindness on the part of the police and the prosecution to the warnings signs in the confession; (3) suppression of and failure to pursue contradictory scientific evidence; and (4) ignorance of the scientific evidence that gave the lie to the alleged confession.

Are there real examples of “flawed forensics” contributing mightily to false convictions? Of course. Do we know how many? Not really, but whatever the precise number may be, there are too many such cases. An article making this now well known point easily could have started with a more a propos example.

Were this the only defect in the article, one might chalk it up to a combination of the expectancy effect and poor research. Perhaps the writer picked the Washington case without worrying too much about the actual facts because he already knew what to expect. (Dare I say that Mr. Stern was not writing on a blank Slate?) Unfortunately, however, there are other inaccuracies in the article. I comment on them in the next posting.

Notes
  1. Mark Joseph Stern, Forensic Science Isn’t Science: Why juries hear—and trust—so much biased, unreliable, inaccurate evidence, Slate, June 11, 2014.
  2. Brandon L. Garrett, Convicting the Innocent: Where Criminal Prosecutions Go Wrong 145 (2011).
  3. Id. at 30.
  4. Id. at 35.
  5. Id. at 147.
  6. Id. at 147-48.
  7. Rather than present a vigorous defense—“[t]he entire defense case lasted only 40 minutes,” id. at 146, Washington’s lawyer—who had never tried a capital case before— “simply asked for the mercy of the jury.” Id. at 154. He did not even point out “the glaring inconsistencies” between the compliant confession and some of the facts in the case—including the race of the white woman who was murdered in front of her two children. Id. at 147. When asked whether she was white or black, Washington chose “black.” Id.
  8. Id. at 148 (note omitted). The “forensic evidence” in question seems to be the following:
    An analyst working for the Virginia Bureau of Forensic Science had tested stains on a central piece of evidence, a blue blanket found on the murdered victim’s bed, and found Transferrin CD, a fairly uncommon plasma protein that is most found in African-Americans. The analyst even ran a second test to double-check the result. The next year, when Earl Washington, Jr., was arrested, they tested his blood and found he did not possess the unusual Transferrin CD. The state did not give the defense the report indicating Washington was excluded by that characteristic. Instead, the state gave the defense an “amended” report. Without having done any new tests, the altered report stated that the results of the Transferrin CD testing “were inconclusive.” The original lab report came to light decades later when Washington filed a civil rights lawsuit after his exoneration.
    Id. at 108 (notes omitted). Inasmuch as the “inconclusive” serum protein test would not have much significance for the defense, I assume that the “rudimentary” blood-typing results that excluded Washington, which the defense saw but overlooked, would have been even more damaging to the prosecution than this amended test for Transferrin CD.

Sunday, 8 June 2014

Kansas Court of Appeals Rejects Post-King Challenge to DNA Collection on Arrest

This year, the Kansas Court of Appeals upheld the state's DNA-on-arrest law against a Fourth Amendment challenge. The result is not surprising in light of the Supreme Court's opinion in Maryland v. King. Still, there are some differences between the Maryland statute and the Kansas one, making the state court's decision not even to publish its opinion a little questionable.

Excerpts from the opinion and two quick comments on them follow:

State v. Biery
No. 109,344, 318 P.3d 1020 (Table)
2014 WL 802100 (Kan. Ct. App. Feb. 28, 2014)

PER CURIAM.

In the early morning hours of May 12, 2012, Hutchinson police observed a white male out walking. The officers approached, without lights or sirens activated or weapons drawn, and asked for identification. Police learned the man was [Willie] Biery and there was an outstanding arrest warrant for his failure to appear for a probation violation hearing. Biery was arrested.
At the jail, Biery emptied his pockets and revealed a small plastic baggie containing methamphetamine. Biery was charged with possession of methamphetamine and booked into the jail for both violations. Because possession of methamphetamine is a felony and his DNA was not on file, Biery was asked to provide a DNA sample, via buccal mouth swab ... . Biery refused. ...

Biery was [found guilty of] refusing to give a DNA sample , in violation of K.S.A.2011 Supp. 21–2511(e)(2). ...

On appeal, Biery's sole issue is whether the statutory scheme for the collection, handling, and storage of DNA samples ... is a violation of the Fourth Amendment to the United States Constitution and § 15 of the Kansas Constitution Bill of Rights. ...
The recent United States Supreme Court decision in [Maryland v.] King, 133 S.Ct. 1958 [2013)], addressed this issue ... As part of their standard procedure for a person arrested and charged with felony offenses, Maryland police took a DNA sample by buccal swab ...

In determining whether the warrantless search was reasonable, the United States Supreme Court held the DNA collection statute served a legitimate government interest by providing a safe and accurate way to process and identify the persons taken into custody, reducing risk to police and those in police custody, ensuring criminals are available to be tried, assessing the danger an individual might pose to the public before setting bond, and reducing the possibility of innocent persons being wrongfully held. ... The Court also noted DNA collection is a search incident to a lawful arrest which, even lacking individual suspicion, is virtually unchallenged in American jurisprudence. ...
Hmm, the Supreme Court did not uphold the DNA collection in Maryland because it fell within the "search incident to arrest" exception to the warrant requirement. That exception only allows to police to search a person and his immediate surroundings to prevent the individual taken into custody from using a weapon or destroying evidence. King applied a balancing test to recognize what is effectually a new exception.

Because the search was minimally intrusive, served a legitimate government interest, was reasonable due to an arrestee's reduced expectation of privacy, and protected against unwarranted disclosures, the Court affirmed the constitutionality of the Maryland statute. ...

... Biery claims the Maryland statute at issue in King is significantly different than the one in place in Kansas ... . Thus, ... Biery argues the Maryland statutory scheme was deemed constitutional because it provided sufficient safeguards against the accidental disclosure or misuse of such samples. Biery claims the Kansas statute lacks such safeguards and, therefore, fails to pass constitutional muster. ...

... Because ... the Kansas Bureau of Investigation (KBI) [must] comply with national standards regarding the collection and maintenance of DNA records, the State argues K.S.A.2011 Supp. 21–2511 provides sufficient statutory safeguards to be considered constitutional under King. ...

The State is correct. ... In regards to the dissemination of DNA information, Kansas law only allows release of DNA records and samples to “authorized criminal justice agencies.” ... Finally, the overall process is governed by the KBI, which “shall promulgate rules and regulations” for the collection and maintenance of samples; expungement and destruction of samples; and procedures in compliance with national standards for DNA records. ...
The Kansas statute differs in a couple of ways that the court does not mention. It applies to a broader class of crimes than the Maryland law. It does not defer the DNA profiling until after an arraignment. It does not require the destruction of samples and profiles if there is no conviction. Apparently, the Kansas court did not consider these differences significant, and it proceeds to present the weaker Kansas provision for expungement as an argument for the reasonableness of the Kansas law.

We pause to note the charges leading to Biery's felony arrest have since been dismissed following the suppression of the evidence against him. ... With the dismissal of his case, K.S.A.2011 Supp. 21–2511(j)(1)(B) provides ... for ... a procedure which allows the defendant to petition to expunge and destroy the DNA samples and profile record in the event of a dismissal of charges, expungement or acquittal at trial.

Had Biery provided a sample, he could now proceed to ask the KBI to expunge and destroy his DNA sample. ...

Biery was lawfully under arrest for a felony at the time the buccal swab was requested. K.S.A.2011 Supp. 21–2511(e) requires that any person subject to a valid felony arrest to submit a buccal swab. The statute does not violate the Fourth Amendment to the United States Constitution or § 15 of the Kansas Constitution Bill of Rights and is constitutional. ...

Friday, 6 June 2014

Fingerprinting Errors and a Scandal in St. Paul

Reports of crime laboratory scandals and fraud have become legion. It is scandalous to convict and imprison people on the basis of contrived, fabricated, or even incompetently generated laboratory evidence. But not all scandals are equal. Consider, for example, the report last year in the ABA Journal that
the St. Paul, Minn., police department’s crime lab suspended its drug analysis and fingerprint examination operations after two assistant public defenders raised serious concerns about the reliability of its testing practices. A subsequent review by two independent consultants identified major flaws in nearly every aspect of the lab’s operation, including dirty equipment, a lack of standard operating procedures, faulty testing techniques, illegible reports, and a woeful ignorance of basic scientific principles. 1/
The article proceeds to describe deplorable conditions in the drug testing lab, but all it says about latent print work is that “[t]he city has since hired a certified fingerprint examiner to run the lab, who has announced plans to resume its fingerprint examination and crime scene processing operations, and begin the procedure for seeking accreditation.”

Curious as to what the latent-print examiners had been doing, I turned to a local newspaper article entititled "St. Paul Crime Lab Errors Rampant." It reported that “[t]he police department hired two consultants to work on improving the lab after a ... Court hearing last year disclosed flawed drug-testing practices” and that “the lab recently resumed fingerprint work by certified analysts.” 2/

One consultant, “Schwarz Forensic Enterprises of Ankeny, Iowa, ... studied the crime lab's latent fingerprint comparison, processing and crime scene units” and found that “[p]ersonnel appeared to have attended seminars and training, but there wasn't formal competency testing or a program to assess ongoing proficiency.”

These untested personnel offered an opportunity to see how poorly monitored analysts performed. Would they succumb to the widely advertised cognitive biases that might cause latent print examiners to declare matches that do not exist? Would they declare matches more frequently than certified examiners? Apparently not:
"'Despite these deficiencies, no evidence of erroneous identifications by latent print examiners was found; but we did find numerous examples of cases wherein examiners had failed to claim latent prints as suitable for identification and/or to identify prints to suspects,'
the Schwarz report said." In other words, the incidence of false negatives and missed opportunities to make identifications or exclusions was high, but no false-positive errors were found. “A review of 246 fingerprint cases found the unit successfully identified prints only ‘in cases where the print detail is of extraordinarily high quality.’”

This outcome is consistent with more rigorous studies showing that when latent print examiners make mistaken comparisons, the errors are usually false exclusions—not false matches. 3/ This tendency reflects a different sort of bias—an unwillingness to declare a match unless the match seems quite clear.

Of course, 246 instances without false positives from worrisome fingerprint analysts does not prove that they never make false matches. If this group were making false identifications 1% of the time, for instance, the probability that no false positives would be seen in a run of 246 independent cases (each with the same 1% false-match probability) would be (1 – .01)246 = 8%.

The absence of false positives also is consistent with an intriguing 2006 report by Itiel Dror and David Charlton. 4/ These investigators had six experienced, certified, and proficiency-tested analysts examine sets of prints from four cases in which, years ago, the examiners had found exclusions and another four cases in which they had made identifications. The subjects did not realize that they had seen these prints before. In some instances of previous exclusions, the examiners were told that a suspect had confessed. In none of these cases did the examiners depart from their earlier judgment of a match.

On the other hand, in cases of previous identifications, when examiners were told that the suspect was in police custody at the time of the crime, two examiners switched from an exclusion to an identification, and one switched to “cannot decide.” Although these sample sizes are too small to justify strong and widely generalizable conclusions, it looks like it is easier for information that is not needed for the analysis to prompt an exclusion than an individualization.

Dror and Charlton interpret their results as supporting (among other things) the claim “that the threshold to make a decision of exclusion is lower than that to make a decision of individualization.” This higher threshold would make it more difficult to bias an examiner to make a false identification than to make a false exclusion.

Did any of the 246 St. Paul cases involve contextual bias of one kind or another? If so, it would be interesting to find out if these examiners resisted contextual suggestions favoring identifications or exclusions in those cases. Audits like these could be helpful not only in getting laboratories with problems back on track, as in St. Paul, but also as a source of information on the risks of different types of errors in various settings and circumstances.

Notes
  1. Mark Hansen, Crime Labs Under the Microscope after a String of Shoddy, Suspect and Fraudulent Results, ABAJ, Sept. 2013
  2. Mara H. Gottfried & Emily Gurnon, St. Paul Crime Lab Errors Rampant, Reviews Find, Pioneer Press, Feb. 14, 2013
  3. See, e.g., Fingerprinting Under the Microscope: Error Rates and Predictive Value, Forensic Science, Statistics, and the Law, April 30, 2012; Fingerprinting Error Rates Down Under, June 24, 2012, Forensic Science, Statistics, and the Law.
  4. Itiel E. Dror & David Charlton, Why Experts Make Errors, 56 J. Forensic Identification 600-16 (2006)

Wednesday, 4 June 2014

Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 3)

After explaining that Florida’s statutory cutoff of –2σx corresponds to an IQ score of 70 (because IQ tests are normed to have a mean of 100 and a standard deviation of 15), Justice Kennedy observes that:
Florida's rule disregards established medical practice in two interrelated ways. It takes an IQ score as final and conclusive evidence of a defendant's intellectual capacity, when experts in the field would consider other evidence. It also relies on a purportedly scientific measurement of the defendant's abilities, his IQ score, while refusing to recognize that the score is, on its own terms, imprecise.
Here, I show that these two limitations on IQ scores are less “interrelated” than Justice Kennedy suggests.

The first issue: validity

The first limitation involves what social scientists call “validity”—the extent to which something measures the real quantity of interest. For example, measuring the volume of a box by attending only to one dimension, such as height, is invalid because it ignores the two other determinative variables of width and depth.

The second limitation concerns the precision or “reliabilility” of the measurement—regardless of validity. Being able to measure the height of a box to the nearest millimeter time after time achieves precision and reliability, but it still lacks validity (with respect to the variable of volume). Moreover, measuring width or depth—even somewhat imprecisely—will add to validity but will do nothing to enhance the precision of the measurement of height.

Likewise, the “other evidence” to which the Court referred does not make IQ scores any more precise. Rather, it relates to what is known in the trade as “adaptive functioning.” The opinion defines adaptive functioning as “the inability to learn basic skills and adjust behavior to changing circumstances.” The Court disparages the “mandatory cutoff” of –2σx because this cut-score means that
sentencing courts cannot consider even substantial and weighty evidence of intellectual disability as measured and made manifest by the defendant's failure or inability to adapt to his social and cultural environment, including medical histories, behavioral records, school tests and reports, and testimony regarding past behavior and family circumstances. This is so even though the medical community accepts that all of this evidence can be probative of intellectual disability, including for individuals who have an IQ test score above 70.
If one were to follow this we-need-another-variable theory of “intellectual disability” to its logical limit, no IQ score could preclude the more comprehensive assay of all forms of “substantial and weighty evidence of intellectual disability.” An individual with above-average IQ scores also might “manifest [a] failure or inability to adapt to his social and cultural environment [as shown by] medical histories, behavioral records, school tests and reports, and testimony regarding past behavior and family circumstances.”

But surely the Court cannot claim that the execution of intellectually gifted but maladapted criminals is cruel and unusual while the execution of intellectually gifted and socially well adjusted criminals is not. To avoid such anomalies, the Court follows the contemporary (and prior) mental health practice of limiting “intellectual disability” to “concurrent deficits in intellectual and adaptive functioning” (emphasis added), which requires “significantly subaverage intellectual functioning” in addition to mere “deficits in adaptive functioning.” If it is clear that an individual is able to function intellectually within a broad but “normal” range, then a state need not entertain a claim of “intellectual disability” based solely on problems in adaptive functioning. Therefore, an IQ within the normal range should suffice to displace the offender from the potentially death-disqualified group.

So the possibility of weighty evidence of deficits in adaptive functioning, although relevant to clinicians, turns out to be no explanation for why Florida cannot draw the line at –2σx. If evidence of an adaptive-function deficit does not bar the state from executing criminals with IQs in the broad range of normalcy, why cannot the state define all IQs above 70 (–2σx) as lying within that range? That “experts in the field would consider other evidence” than IQ scores is not an answer. The answer has to be that (1) there is a range above which IQ, in of of itself, is a valid measure of the absence of “intellectual disability,” and (2) this range does not extend all the way down to 70. When these conditions hold, experts would not (or need not) consider “other evidence.”

Ironically, the Court’s opinion contradicts the second proposition. It clearly implies that the state could use a perfectly precise IQ measurement just above –2σx as conclusive evidence of intellectual disability. But if that is so, then the problem is not the failure to allow evidence of adaptive functioning. It is solely the existence of nonzero measurement error of IQ alone.

The second issue: precision (reliability)

Apparently (and dubiously) reserving the term “scientific” for precise measurements, Justice Kennedy stated that the “purportedly scientific measurement of the defendant's abilities, his IQ score, ... is, on its own terms, imprecise.” The problem is that although “there is evidence that Florida's Legislature intended to include the measurement error in the calculation ... the Florida Supreme Court ... has held that a person whose test score is above 70, including a score within the margin for measurement error, does not have an intellectual disability ... .”

In other words, a legislature that wants to preclude the more elaborate evaluations of all offenders with IQ scores below 70 could do so if only it had a way to measure IQs with perfect accuracy. Because of the “measurement error” of IQ tests, this legislature must adopt a higher cutoff. The Court, relying on the diagnostic literature, repeatedly refers to a cutoff of 75 as assuring an adequate safety margin.

The dissent had harsh words for the choice of 75, and I will get to those later, after examining where the figure of 75 comes from. At this point, no excursion into statistical theory is required to recognize that there is something weird about saying that IQ scores are problematic because they are an incomplete measure of “intellectual disability,” but then using them—and only them—within a band that accounts only for the error in measuring IQ. By definition, this band does not attend to the other factors that should be part of the full analysis. To put it another way, if the problem lies with using IQ alone, the solution lies in defining the range of IQ scores in which the other factors realistically could produce a different diagnosis. However, the error in IQ measurements has no clear connection to the range in which the failure to look beyond IQ makes a difference.

The majority’s response is essentially that if the mental health profession generally agrees that incompleteness is only a significant concern within the logically unrelated range of IQ-score error, then that is all that the Cruel and Unusual Punishment Clause demands. To which the dissent replies that abdicating the line drawing to the professionals makes no constitutional sense and “will also lead to serious practical problems.”

The dissent’s peculiar proof of “instability”

The first such problem is “instability.” According to Justice Alito:
This danger is dramatically illustrated by the most recent publication of the APA, on which the Court relies. This publication fundamentally alters the first prong of the longstanding, two-pronged definition of intellectual disability that was embraced by Atkins and has been adopted by most States. In this new publication, the APA discards “significantly subaverage intellectual functioning” as an element of the intellectual-disability test. Elevating the APA's current views to constitutional significance therefore throws into question the basic approach that Atkins approved and that most of the States have followed. 1/
The American Psychiatric Association’s latest version of its venerable Diagnostic and Statistical Manual of Mental Disorders—the DSM-5—“was published in May 2013 amid a storm of controversy and bitter criticism.” 2/ In general, critics maintain that “D.S.M.’s diagnostic categories lacked validity, that they were not ‘based on any objective measures,’ and that, ‘unlike our definitions of ischemic heart disease, lymphoma or AIDS,’ which are grounded in biology, they were nothing more than constructs put together by committees of experts.” 3/ Neither opinion even hints at such turmoil. The majority genuflects to clinical expertise and guidelines. The dissent raises no questions about validity and subjectivity, but objects to substituting “the standards of professional associations, which at best represent the views of a small professional elite” for “the standards of the American people.”

As for “instability,” the DSM-5 has brought a profusion of new or redefined disorders, but it does not radically change the definition of “intellectual disability” or dispense with the criterion of “significantly subaverage intellectual functioning.” It simply substitutes the word “intellectual ... deficit” for “significantly subaverage.” The diagnostic criteria have remained remarkably similar over the 19 years between the DSM-4 and the DSM-5.

The DSM-5 specifies that “[t]he first diagnostic criterion that “must be met” is “A. Deficits in intellectual functions ... confirmed by ... both clinical assessment and standardized intelligence testing.” If the intelligence testing does not demonstrate subaverage performance, it is hard to see how it could confirm the existence of a meaningful deficit. Moreover, the DSM-5 elaborates, making it plain that significantly subaverage IQ remains a sine qua non for the diagnosis:
The essential features ... are deficits in general mental abilities (Criterion A) and impairment in everyday adaptive functioning ... (Criterion B) [with o]nset is during the developmental period (Criterion C). The diagnosis of ... is based on both clinical assessment and standardized testing ... . Intellectual functioning is typically measured with ... tests of intelligence. Individuals with intellectual disability have scores of approximately two standard deviations or more below the population mean, including a margin for measurement error (generally +5 points). On tests with a standard deviation of 15 and a mean of 100, this involves a score of 65–75 (70 ± 5).
Compare this to the DSM-4 (or the DSM-4-TR cited by Justice Alito, which uses the same words):
The essential feature of Mental Retardation is significantly subaverage general intellectual functioning (Criterion A) that is accompanied by significant limitations in adaptive functioning ... (Criterion B) [with] onset ... before age 18 years (Criterion C). ... General intellectual functioning is defined by the intelligence quotient ... obtained by assessment with ... intelligence tests ... . Significantly subaverage intellectual functioning is defined as an IQ of about 70 or below (approximately 2 standard deviations below the mean). It should be noted that there is a measurement error of approximately 5 points in assessing IQ, although this may vary from instrument to instrument ... . Thus, it is possible to diagnose Mental Retardation in individuals with IQs between 70 and 75 who exhibit significant deficits in adaptive behavior. Conversely, Mental Retardation would not be diagnosed in an individual with an IQ lower than 70 if there are no significant deficits or impairments in adaptive functioning.
Thus, there are wording changes over the 19 years from 1994 to 2013, but Criterion A remains Criterion A, IQ tests remain critical to the diagnosis, and the range of test scores that lend themselves to the diagnosis is the same. The APA has changed the emphasis somewhat, and it has spelled out the constructs a little more (in words not quoted here). Nevertheless, to claim that the shift “dramatically illustrate[s a] fundamental[] alter[ation in] ... the longstanding ... definition of intellectual disability” seems, well, melodramatic.

State laws that rely on –2σx plus a margin of safety for measurement error, are compatible with Atkins, Hall, DSM-4, and DSM-5. Of course, whether this is a logically or functionally appropriate manner of defining “intellectual disability” for purposes of capital punishment is open to debate. Resolving this debate requires a more detailed and accurate understanding of the concept of measurement error than the Hall opinions provide.

Footnotes
  1. The second problem is that “changes adopted by professional associations are sometimes rescinded.” This problem is just a form of instability. The third problem is hypothetical (thus far) as it relates to intellectual disability determinations: “what if professional organizations disagree? The Court provides no guidance for deciding which organizations' views should govern.” The fourth and final “practical problem” is actually conceptual—and quite important. “[D]efinitions of intellectual disability ... are promulgated for use in making a variety of decisions that are quite different from the decision whether the imposition of a death sentence in a particular case would serve a valid penological end. ... [I]n determining eligibility for social services, adaptive functioning may be much more important.”
  2. Nat’l Health Service Choices, Controversy over DSM-5: New Mental Health Guide, Aug. 15, 2013.
  3. Gary Greenberg, The Rats of N.I.M.H., New Yorker, May 16, 2013 (quoting Thomas Insel, the director of the National Institute of Mental Health). 

Other postings in this series
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 1) (introduction)
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2) (on standard deviation)

Monday, 2 June 2014

Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2)

Justice Kennedy described the Florida law that prompted the trial court to reject Hall’s claim of intellectual disability as follows:
Florida's statute defines intellectual disability for purposes of an Atkins proceeding as “significantly subaverage general intellectual functioning existing concurrently with deficits in adaptive behavior and manifested during the period from conception to age 18.” Fla. Stat. § 921.137(1) (2013). The statute further defines “significantly subaverage general intellectual functioning” as “performance that is two or more standard deviations from the mean score on a standardized intelligence test.” Ibid. The mean IQ test score is 100. The concept of standard deviation describes how scores are dispersed in a population. Standard deviation is distinct from standard error of measurement, a concept which describes the reliability of a test and is discussed further below. The standard deviation on an IQ test is approximately 15 points, and so two standard deviations is approximately 30 points. Thus a test taker who performs “two or more standard deviations from the mean” will score approximately 30 points below the mean on an IQ test, i.e., a score of approximately 70 points.
Because standard deviations are fundamental to the Florida law, to the Court’s conclusions about it, and to its dicta regarding the lowest mandatory IQ cut-off that a state can use, I am going to be persnickety in unpacking this paragraph.

Although the Court is only discussing the standard deviation of IQ scores, “standard deviation” (SD) has a much broader meaning. When one considers the general meaning of the term, it becomes clear that the SD of the scores does not quite describe “how scores are dispersed in a population.” It merely indicates how much they are dispersed in either a population or a sample.

For example, the trial court heard testimony about at least four IQ test scores for Hall—71, 72, 73, and 80. (It declined to consider a score of 69 on another test because the psychologist who administered and scored the test was dead and Hall’s counsel had violated an order “to provide the State with the [underlying] testing materials and raw data.” Indeed, Hall had taken as many as nine IQ tests over a 40-year period.)  The standard deviation of the scores considered by the trial court is the square root of the average squared deviation from the mean—namely,

SD = {[(71–74)2 + (72–74)2 + (73–74)2 + (80–74)2]/4}1/2 = 3.53.

Tossing in the excluded score of 69 increases the SD to 3.74. The SD increases because the additional score is below the range of the other four, thus creating more variability in the sample (and lowering the mean from 74 to 73).

Of course, the Court’s number of 15 for the SD of IQ scores does not come from Hall’s scores. At this point, I use his scores only to elaborate on the Court’s observation that a standard deviation is a statistic that indicates how much the numbers in some set of numbers fluctuate around their mean. The standard deviation of 15 IQ points is an estimate of how much the scores of everyone in the general population—a large batch of numbers indeed—would vary if everyone took the test. The average score would be approximately 100, and there would be a lot of scatter around this mean. (In fact, the raw scores on the test are transformed in light of their mean and SD to force them to have a desired mean and SD near 100 and 15, respectively.) And, yes, 100 – (2×15) = 70, so Florida’s choice of 2 SDs to demarcate “significantly subaverage general intellectual functioning” translates into a score of 70 on a test with this mean and SD.

But the fact that every batch of numbers has a SD does not tell us “how [these numbers] are dispersed.” The numbers could be highly concentrated around a single value, with outliers on the flanks. Their distribution could be flat, with an equal fraction of the numbers spread out everywhere. The distribution might show clustering at several locations, and so on.

IQ scores, however, are dispersed approximately according to a “normal” or “Gaussian” curve. This distribution is the bell-shaped one prominent in elementary statistics courses. There are other bell-shaped curves, and all kinds of other interesting and important families of curves, but IQ scores, like many physical variables (such as weight and height), tend to be normally distributed across the members of a population (and hence in representative samples of that population).

The exact shape of all such normal distributions can be determined from two numbers—the mean and the standard deviation. The mean states where the bell sits, and the standard deviation determines how steeply its sides flow down from the top.You can see for yourself by entering your favorite means and standard deviations into the demonstration program in the OnlineStatBook.

Using the variable X to denote IQ scores and the symbol σx to designate their standard deviation, the particular normal distribution used in the Court’s calculation is such that, 2.28% of the scores lie below 70 (which, as the Court calculated it, corresponds to –2σx), and 4.75% fall below 75 (which, for the mean of 100 and standard deviation σx of 15, corresponds to –1.67σx). The latter IQ score, x = 75, is significant because Hall conceded (and the Court seemed to agree) that Florida could have chosen this score as its cut-off. For example, the Court expressed dissatisfaction that, in light of its calculations, the effect of Florida’s cut-off of –2σx was to preclude legally effective “professional[] diagnose[s of] intellectual disability [in a case like Hall’s, for which] the individual's IQ score is 75 or below.”

The dissent insisted that states should have more discretion to set cut-off scores. Unless –1.67σx (or 75) corresponds to the level of impairment that justifies a categorical rule, the majority has no satisfying reason to select one cut-off over the other. Why is the Court’s choice of 1.67 standard deviations below the mean the highest that the Constitution permits? Why is Florida’s two-standard-deviation rule insufficient?

The Court’s answer leans heavily on the standard error of measurement — another technical term that appears in the paragraph quoted above: “Standard deviation is distinct from standard error of measurement, a concept which describes the reliability of a test and is discussed further below.” But the standard error of measurement is also a standard deviation, one that is estimated, almost magically, from test reliability statistics. Thus, a more precise sentence would have been: “The standard deviation of all test scores is distinct from another standard deviation known as the standard error of measurement, which depends on the reliability of the test. We discuss the standard error of measurement below.”

Other postings in this series

  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 1), May 29, 2014 (introduction)
  •  Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2), June 2, 2014 (on standard deviation)
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 3), June 4, 2014 (on validity and the stability of the APA's diagnostic criteria)

Saturday, 31 May 2014

Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 1)

In Hall v. Florida, 134 S.Ct. 1986 (2014), the Supreme Court struck down Florida’s practice of using an IQ score of 70 points or less as a dispositive measure of the level of the “intellectual disability” that precludes capital punishment. Justice Kennedy’s opinion (joined by Justices Breyer, Ginsburg, Sotomayor, and Kagan) summarizes the case pithily:
This Court has held that the Eighth and Fourteenth Amendments to the Constitution forbid the execution of persons with intellectual disability. Atkins v. Virginia, 536 U.S. 304, 321 (2002). Florida law defines intellectual disability to require an IQ test score of 70 or less. If, from test scores, a prisoner is deemed to have an IQ above 70, all further exploration of intellectual disability is foreclosed. This rigid rule, the Court now holds, creates an unacceptable risk that persons with intellectual disability will be executed, and thus is unconstitutional.
Id. at 1990. Led by Justice Alito, Chief Justice Roberts and Justices Scalia, Kennedy, and Thomas dissented. The Court, they maintained, was overruling Atkins and adopting “a uniform national rule that is both conceptually unsound and likely to result in confusion.” Id. at 2002 (Alito, J., dissenting). Among other things, the dissent warns that the Court “misunderstands” the statistical concepts of standard error and confidence intervals, id. at 2009, and that it therefore “makes factual mistakes that will surely confuse States attempting to comply with its opinion.” Id. at 2010.

The dissent has a point. Parts of the majority opinion are elliptical and potentially confusing. Nonetheless, some of the harsh critique is overdrawn. Moreover, Justice Alito's presentation of psychometric concepts also is hardly error-free, inviting a rejoinder of "tu quoque."

Therefore, in a series of postings yet to come, I will, in Justice Alito's words, "wade[] into technical matters that must be understood in order to see where the Court goes wrong." But I'll do the same for the dissenting opinion's presentation of these matters.

Other postings in this series
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 2), June 2, 2014 (on standard deviation)
  • Quarreling and Quibbling over Psychometrics in Hall v. Florida (part 3), June 4, 2014 (on validity and the stability of the APA's diagnostic criteria)

Saturday, 10 May 2014

Latent Fingerprints and the Uniqueness Hypothesis

It has been argued many times that assertions of the uniqueness of all fingerprints are insufficient to warrant claims that a latent print must match the fingerprints of one and only one individual in the world. However, some fingerprint analysts and criminalists, remaining true to the faith of an earlier generation (e.g., Swofford 2012; Vanderkolk 2012), continue to rely on uniqueness as the conceptual and practical foundation for their work.

In this regard, it may be worth noting a paper from two computer scientists that estimates the probability that the latent print that the FBI misidentified in the Madrid train bombing case back in 2004 would have matched at least one individual in the AFIS database. According to the abstract of Su and Shrihari (2010a):
While tremendous efforts have been made in 10-print individuality studies, latent fingerprint rarity continues to be a difficult problem and has never been solved because of the small finger area and poor impression quality. The proposed method is able to predict the core points of latent prints using Gaussian processes and align the latent prints by overlapping the core points. A novel generative model is also proposed to take into account the dependency on nearby minutiae and the confidence of minutiae in the probability of random correspondence calculation. The new methods are illustrated by experiments on the well-known Madrid bombing case. The results show that the probability that at least one fingerprint in the FBI IAFIS databases (over 470 million fingerprints) matches the bomb site latent is 0.93 which is large enough to lead to misidentification.
To be clear, Su and Shrihari do not suggest that full fingerprints of anyone in the AFIS database match those of the Madrid bomber. Their point is that the features of the latent prints that led to the misidentification of Mayfield are not likely to be unique. But surely, when it comes to thinking about the relevance of the uniqueness hypothesis for fingerprints and assessing the probative value of latent fingerprint identification, this is what matters.

The paper overlaps another one (Su and Shrihari 2010b) by the same authors delivered at another conference. I have not searched for responses to their work.

References

Chang Su & Sargur N. Srihari, Latent Fingerprint Rarity Analysis in Madrid Bombing Case, in Computational Forensics: 4th International Workshop, IWCF 2010, Tokyo, Japan, November 11-12, 2010a, Revised Selected Papers (Sako, Hiroshi; Franke, Katrin; Saitoh, Shuji eds. 2011), Lecture Notes in Computer Science, 6540:173-184

Chang Su & Sargur Srihari, Evaluation of Rarity of Fingerprints in Forensics, in Proceedings of Neural Information Processing Systems, Vancouver, Canada, Dec. 6-9, 2010b

Henry J. Swofford, Individualization Using Friction Skin Impressions: Scientifically Reliable, Legally Valid, J Forensic Identification 62:65-79, 2012

John R. Vanderkolk, Examination Process, in The Fingerprint Sourcebook 9-3 to 9-26, 2012