Showing posts with label Daubert. Show all posts
Showing posts with label Daubert. Show all posts

Thursday, 12 June 2014

More on the Mistakes in “Forensic Science Isn’t Science”

In "Flawed Journalism on Flawed Forensics in Slate Magazine" I referred to "other inaccuracies in the article" by Mark Joseph Stern entitled “Forensic Isn’t Science.” Here are some excerpts from the article with some thoughts.

Behind the myriad technical defects of modern forensics lie two extremely basic scientific problems. The first is a pretty clear case of cognitive bias: A startling number of forensics analysts are told by prosecutors what they think the result of any given test will be. This isn’t mere prosecutorial mischief; analysts often ask for as much information about the case as possible—including the identity of the suspect—claiming it helps them know what to look for. Even the most upright analyst is liable to be subconsciously swayed when she already has a conclusion in mind. Yet few forensics labs follow the typical blind experiment model to eliminate bias. Instead, they reenact a small-scale version of Inception, in which analysts are unconsciously convinced of their conclusion before their experiment even begins.

Few psychologists would propose that expectancy effects will cause errors in interpretation in every experiment. The role of cognitive bias in science and inference generally is quite complicated. There is no typical “blind experiment model.” The physicists who announced the discovery of the Higgs boson or those who believed they detected the marks of gravitational waves in the cosmic background radiation had no such model. In much of science, experimenters know what they are looking for; fortunately, some results are not ambiguous and not subject to subtle misinterpretation. There is the joke that if your experiment needs statistics, you should do a better experiment.

That said, when interpretations are more malleable—as is often the case in many forensic disciplines—various methods are available to minimize the chance of this source of error. One is “sequential unmasking” to protect forensic analysts from unconscious (or conscious) bias that could lead them to misconstrue their data when exposed to information that they do not need to know. There is rarely, if ever, an excuse for not using methods like these. But their absence does not make a solid latent fingerprint match or a matching pair of clean, single-source electropherograms, for example, into small-scale versions of Inception.

Without a government agency overseeing the field, forensic analysts had no incentive to subject their tests to stricter scrutiny. Groups such as the Innocence Project have continually put pressure on the Department of Justice—which almost certainly should have supervised crime labs from the start—to regulate forensics. But until recently, no agency has been willing to wade into the decentralized mess that hundreds of labs across the country had unintentionally created.

When and where was the start of forensic science? Ancient China? Renaissance Europe? Should the U.S. Department of Justice have been supervising the Los Angeles Police Department when it founded the first crime laboratory in the United States in 1923? The DOJ has had enough trouble with the FBI laboratory, whose blunders led to reports from the DOJ’s Office of the Inspector General and which has turned to the National Research Council for advice on more than one occasion. The 2009 report of a committee of the National Research Council had much better ideas for improving the practice of forensic science in the United States. Its recommendation of a new agency entirely outside of the Department of Justice for setting standards and funding research, however, gained little political traction. The current National Commission on Forensic Science is a distorted, toothless, and temporary version of the idea.

In 2009, a National Academy of Sciences committee embarked on a long-overdue quest to study typical forensics analyses with an appropriate level of scientific scrutiny—and the results were deeply chilling.

The committee did not undertake a “scientific” study. It engaged in a policy-oriented review of the state of forensic science without applying any particularly scientific methods. (This is not a criticism of the committee. NRC committees generally collect and review relevant literature and views rather than undertake scientific research of their own.) This committee's quest did not begin in 2009. That is when it ended. Congress voted to fund the study in 2005.

Aside from DNA analysis, not a single forensic practice held up to rigorous inspection. The committee condemned common methods of fingerprint and hair analysis, questioning their accuracy, consistent application, and general validity. Bite-mark analysis—frequently employed in rape and murder cases, including capital cases—was subject to special scorn; the committee questioned whether bite marks could ever be used to positively identify a perpetrator. Ballistics and handwriting analysis, the committee noted, are also based on tenuous and largely untested science.

The report is far more nuanced (some might say conflicted) than this. Here are some excerpts:
"The chemical foundations for the analysis of controlled substances are sound, and there exists an adequate understanding of the uncertainties and potential errors. SWGDRUG has established a fairly complete set of recommended practices." P. 135.

"Historically, friction ridge analysis has served as a valuable tool, both to identify the guilty and to exclude the innocent. Because of the amount of detail available in friction ridges, it seems plausible that a careful comparison of two impressions can accurately discern whether or not they had a common source. Although there is limited information about the accuracy and reliability of friction ridge analyses, claims that these analyses have zero error rates are not scientifically plausible." P. 142.

"Toolmark and firearms analysis suffers from the same limitations discussed above for impression evidence. Because not enough is known about the variabilities among individual tools and guns, we are not able to specify how many points of similarity are necessary for a given level of confidence in the result. Sufficient studies have not been done to understand the reliability and repeatability of the methods. The committee agrees that class characteristics are helpful in narrowing the pool of tools that may have left a distinctive mark. Individual patterns from manufacture or from wear might, in some cases, be distinctive enough to suggest one particular source, but additional studies should be performed to make the process of individualization more precise and repeatable." P. 154.

"Forensic hair examiners generally recognize that various physical characteristics of hairs can be identified and are sufficiently different among individuals that they can be useful in including, or excluding, certain persons from the pool of possible sources of the hair. The results of analyses from hair comparisons typically are accepted as class associations; that is, a conclusion of a 'match' means only that the hair could have come from any person whose hair exhibited—within some levels of measurement uncertainties—the same microscopic characteristics, but it cannot uniquely identify one person. However, this information might be sufficiently useful to 'narrow the pool' by excluding certain persons as sources of the hair." P. 160.

"The scientific basis for handwriting comparisons needs to be strengthened. Recent studies have increased our understanding of the individuality and consistency of handwriting and computer studies and suggest that there may be a scientific basis for handwriting comparison, at least in the absence of intentional obfuscation or forgery. Although there has been only limited research to quantify the reliability and replicability of the practices used by trained document examiners, the committee agrees that there may be some value in handwriting analysis."

"Analysis of inks and paper, being based on well-understood chemistry, presumably rests on a firmer scientific foundation. However, the committee did not receive input on these fairly specialized methods and cannot offer a definitive view regarding the soundness of these methods or of their execution in practice." Pp. 166-67

"As is the case with fiber evidence, analysis of paints and coatings is based on a solid foundation of chemistry to enable class identification." P. 170

"The scientific foundations exist to support the analysis of explosions, because such analysis is based primarily on well-established chemistry." P. 172

"Despite the inherent weaknesses involved in bite mark comparison, it is reasonable to assume that the process can sometimes reliably exclude suspects. Although the methods of collection of bite mark evidence are relatively noncontroversial, there is considerable dispute about the value and reliability of the collected data for interpretation." P. 176

"Scientific studies support some aspects of bloodstain pattern analysis. One can tell, for example, if the blood spattered quickly or slowly, but some experts extrapolate far beyond what can be supported." P. 178
Hardly a ringing endorsement of all police lab techniques, but neither is the report an outright rejection of all or even most techniques now in use.

The report amounted to a searing condemnation of the current practice of forensics and an ominous warning that death row may be filled with innocents.

According to an NRC press release issued in February, 2009, “[t]he report offers no judgment about past convictions or pending cases, and it offers no view as to whether the courts should reassess cases that already have been tried.” Such language may be a compromise among the disparate committee members. But to derive the conclusion that death row is “filled with innocents” even partly from the actual contents of the report, one would have to consider the deficiencies identified in the system, the extent to which these deficiencies generated the evidence used in capital cases, and the other evidence in those cases. Other research is far more helpful in evaluating the prevalence of false convictions.

Given the flimsy foundation upon which the field of forensics is based, you might wonder why judges still allow it into the courtroom.

As the 2009 NRC Committee explained, there is no single "field of forensics." Rather,
"Wide variability exists across forensic science disciplines with regard to techniques, methodologies, reliability, error rates, reporting, underlying research, general acceptability, and the educational background of its practitioners. Some of the forensic science disciplines are laboratory based (e.g., nuclear and mitochondrial DNA analysis, toxicology, and drug analysis); others are based on expert interpretation of observed patterns (e.g., fingerprints, writing samples, toolmarks, bite marks, and specimens such as fibers, hair, and fire debris). Some methods result in class evidence and some in the identification of a specific individual—with the associated uncertainties. The level of scientific development and evaluation varies substantially among the forensic science disciplines." P. 182.
The courts have been lax in responding to overblown testimony in some fields and to those techniques that lack proof of their fundamental precepts.

In 1993, the Supreme Court announced a new test, dubbed the "Daubert standard," to help federal judges determine what scientific evidence is reliable enough to be introduced at trial. The Daubert standard ... wound up frustrating judges and scientists alike. As one dissenter griped, the new test essentially turned judges into "amateur scientists," forced to sift through competing theories to determine what is truly scientific and what is not.

Blaming the persistence of the admissibility of the most dubious forensic disciplines on Daubert is strange. Daubert's standard did not spring into existence fully formed, like Athena from the brow of Zeus. A similar standard was in place in a number of jurisdictions. As the The New Wigmore: A Treatise on Evidence shows, the Court borrowed from these cases. A smaller point to note is that there were not one, but two partial dissenters (who concurred in the unanimous judgment). Chief Justice Rehnquist and Justice Stephens objected to the majority’s proffering "general observations" about scientific validity, and they did not complain about the ones the Mr. Stern points to as an explanation for the persistence of questionable forensic "science."

Even more puzzlingly, the new standards called for judges to ask "whether [the technique] has attracted widespread acceptance within a relevant scientific community"—which, as a frustrated federal judge pointed out, required judges to play referee between "vigorous and sincere disagreements" about "the very cutting edge of scientific research, where fact meets theory and certainty dissolves into probability."

That’s Chief Judge Alex Kozinski of the Ninth Circuit Court of Appeals writing on remand in Daubert itself. Judge Kozinski could not possibly be objecting to the Supreme Court's opinion on the ground that "widespread acceptance within a relevant scientific community" is an impenetrable standard. Quite the opposite. He applied that very standard in his previous opinion in the case and was bemoaning what he called the "brave new world" that the Court ushered in as it vacated his opinion. A recent survey of judges found that 96% of the (disappointingly small) fraction responding deemed the general scientific acceptance to be helpful -- more than any other commonly used factor in judging the validity of scientific evidence.

American jurors today expect a constant parade of forensic evidence during trials. They also refuse to believe that this evidence might ever be faulty. Lawyers call this the CSI effect, after the popular procedural that portrays forensics as the ultimate truth in crime investigation. [¶] “Once a jury hears something scientific, there’s a kind of mythical infallibility to it,” Peter Neufeld, a co-founder of the Innocence Project, told me. “That’s the association when a person in white lab coat takes the witness stand. By that point—once the jury’s heard it—it’s too late to convince them that maybe the science isn’t so infallible.”

Refusal to question scientific evidence is not what most lawyers call the CSI effect. In any event, jury research does not support the idea that jurors inevitably reject attacks on scientific testimony or that the testimony of the first witness in a figurative white coat is unshakeable.

If judges can’t be trusted to keep spurious forensic analysis out of the courtroom, and juries can’t be trusted to disregard it, then how are we going to keep the next Earl Washington off death row? One option would be to permit anybody convicted on the basis of biological evidence to subject that evidence to DNA analysis—which is, after all, the one form of forensics that scientists agree actually works. But in 2009, the Supreme Court ruled that convicts had no such constitutional right, even where they can show a reasonable probability that DNA analysis would prove their innocence. (The ruling was 5–4, with the usual suspects lining up against convicts’ rights.)

This option, which applies applies to a limited set of cases (and hence is no general solution) is not foreclosed by District Attorney for the Third Judicial District v. Osborne, 129 S.Ct. 2308 (2009). If there is a minimally plausible claim of actual innocence after conviction, let’s allow such testing by statute. Of course, it would be better to thoroughly test potential DNA evidence (when it is relevant) before trial—something that Osborne's trial counsel declined to request, fearing that it would only strengthen the prosecution’s case.

Until lab technicians follow some uniform guidelines and abandon the dubious techniques glamorized on shows like CSI, forensic science will barely qualify as a science at all. As a recent investigation by Chemical & Engineering News revealed, little progress has been made in the five years since the National Academy of Sciences condemned modern forensic techniques.

Again, as the NRC committee stressed, "forensic science" is not a single, uniform discipline. Since the report, funding has increased, some guidelines have been revised, and significant research has appeared in some fields. Still, the pace resembles that of global warming. It is coming, notwithstanding resistance described in earlier years on this blog.

As for the "investigation by Chemical & Engineering News," the latest I saw from that publication was an article in a May 12, 2014 issue with a map showing selected instances of examiner misconduct dating back to 1993 and indicating that only five states require laboratory accreditation. No effort was made to ascertain how many labs are still operating without accreditation. With no apparent literature review, The article simply asserted that
[I]n the years since [2009], little has been done to shore up the discipline’s scientific base or to make sure that its methods don’t result in wrongful convictions. Quality standards for forensic laboratories remain inconsistent. And funding to implement improvements is scarce. [¶] While politicians and government workers debate changes that could help, fraudsters like forensic chemist Annie Dookhan keep operating in the system. No reform could stop a criminal intent on doing wrong, but a better system might have shown warning signs sooner. And it likely would have prevented some of the larger, systemic problems at the Massachusetts forensics lab where Dookhan worked.
I must be missing the real investigation that the C&E News writers conducted.

Sunday, 23 June 2013

Mathematics on Appeal: Hamilton’s Equations in Lapsley v. Xtek, Inc.

Better off Ted is a satiric TV series about wacky scientists and managers at the amoral high-tech company, Veridian Dynamics. In Lapsley v. Xtek, Inc., 689 F.3d 802, 812 (7th Cir. 2012), Meridian Engineering supplied the scientific breakthrough -- about dynamics -- for plaintiff. But the resemblance ends with the company names. There is nothing funny about this case.

In Xtek, a machine at a steel rolling mill accidentally ejected industrial grease with such force that it shot through a worker’s body, permanently disabling him. As the court of appeals explained:
At trial the jury found that the accident was caused by a design defect in a heavy industrial product designed and manufactured by defendant Xtek, and sold and installed in the mill. That equipment contained an internal spring that could exert over ten thousand pounds of force. The jury accepted the theory of plaintiffs' expert witness, Dr. Gary Hutter, that the spring was the culprit mechanism behind the accident and that an alternative design of a thrust plate in the equipment would have prevented the disabling accident. Xtek has appealed, challenging the district court's denial of its Daubert motion that sought to bar Dr. Hutter from offering his expert opinions, which were essential to the plaintiffs' case.
Xtek’s Daubert challenge, as described in the opinion, was that Dr. Hutter, an engineer who founded the consulting firm, was performing “the simulation of science, not science,” when he used a mathematical model to infer the cause of the release and to conclude that an alternative design would have prevented it. He did no empirical testing to verify his theories, and to Xtek’s lawyers, the notes attached to Dr. Hutter’s report were the “equivalent of Sanskrit.”

The U.S. Court of Appeals for the Seventh Circuit squashed this argument. The opinion could have been rather mundane. It need only have stated that (1) Dr. Hutter used equations that were not in dispute; (2) defendant raised no question about the expert’s assumptions; and (3) supplementary physical testing was not feasible.

In addition to making these points, however, the court took the opportunity to lecture counsel about science. The lecture began with the advice that
Lawyers and judges who were not trained in science can benefit from the famous “Two Cultures” lecture given in 1959 by British scientist and novelist C.P. Snow, in which he described the cultural gap between persons schooled in the sciences and those schooled in the humanities:
A good many times I have been present at gatherings of people who, by the standards of the traditional culture, are thought highly educated and who have with considerable gusto been expressing their incredulity at the illiteracy of scientists. Once or twice I have been provoked and have asked the company how many of them could describe the Second Law of Thermodynamics. The response was cold: it was also negative. Yet I was asking something which is about the scientific equivalent of: Have you read a work of Shakespeare's?
Nowadays, the opinion admonished, “[j]udges and lawyers do not have the luxury of functional illiteracy in either of these two cultures.” The trial judge was not guilty of such illiteracy. There was “no indication, either from the district court's Daubert ruling or its later discussions of the expert evidence during trial, of any deficiency in the court's preparation or in its understanding of the proposed evidence.”

This leaves the lawyers. “For the curious” the court supplied a link to a digitalized version of Sir Isaac Newton’s Philosophise Naturalis Principia Mathematica. This classic is not written in Sanksrit, but the original Latin is challenging enough. Consequently, the court reproduced and described the equations for force, kinetic energy, and pressure. After a nod to quantum mechanics (but not relativity), it announced that “Newtonian physics still provides a reliable and workable description for the mechanical systems of a steel mill.” Seeing no problem with the expert’s use of these “basic equations of classical mechanics” and invoking Galileo and Descartes in addition to Newton and two of the Bernouillis (Daniel and Johann, whose animosity was such that they would not have appreciated being mentioned in the same sentence), the court concluded that “physical tests of [the] theories with regard to causation (the effect of the spring releasing) and alternate design (the reduction in pressure from the grease grooves)” were not essential.

Despite its unusual historical and mathematical sweep, Xtek has a narrow holding. It does not approve of all mathematical modeling of complex phenomena in lieu of physical tests. To be sure, the opinion asserts that “[a] mathematical or computer model is a perfectly acceptable form of test,” and “that simulation is one of the most common of scientific and engineering tools.” Moreover, it contains the memorable line,“We do not require experts to drop a proverbial apple each time they wish to use Newton's gravitational constant in an equation.” But that analogy only goes so far. The reason physicists need not remeasure G every time they wish to use the inverse-square law of gravitational force is that the value already is known from many experiments. In Xtek, the question is the value of an unknown quantity.

The holding is simply that when the equations (including the constants in them) are familiar and their applicability is unchallenged, an expert can use them to make relevant computations. Additional physical testing may be required under Daubert in some cases, but not when it would be too expensive or dangerous. “Around the world, computers simulate nuclear explosions, quantum mechanical interactions, atmospheric weather patterns, and innumerable other systems that are difficult or impossible to observe directly” (emphasis added). Thus, Dr. Hutter did not have “to try to recreate the binding up of a ten thousand pound spring to produce a potentially deadly jet of industrial grease ... to testify to the results of his mathematical simulations.”

It is not often that one encounters equations in a judicial opinion. Years ago, in Branion v. Gramly, 855 F.2d 1256 (7th Cir. 1988), Judge Frank Easterbrook tossed out a partial differential equation and a few calculations of his own to chastise a habeas corpus petitioner’s lawyers for “fooling with algebra” in their brief. I, 1991, I described that effort as needlessly opaque but basically correct.

In this case, the author of the opinion, Judge David Hamilton, stuck to simpler equations. And that takes me to another similarity in names. Judge Hamilton's equations also are simpler than the elegant formulation of dynamics known to physicists as Hamilton’s equations (invented by William Rowan Hamilton, 1805-1865). 

References
  • David H. Kaye, David E. Bernstein Jennifer L. Mnookin, The New Wigmore: A Treatise on Evidence: Expert Evidence, New York: Aspen Pub. Co., 2d ed., 2011(updated annually)
  • David H. Kaye, Statistics for Lawyers and Law for Statistics, Michigan Law Review, Vol. 89, No. 6, May 1991, pp. 1520-1544 (review essay discussing Branion v. Gramly)

Wednesday, 17 October 2012

More on Semrau: The Other Daubert Factors

In United States v. Semrau, the U.S. Court of Appeals for the Sixth Circuit upheld the exclusion of a defendant’s “unilateral” fMRI testing for conscious deception. Previously, I focused on the court’s discussion of error rates. Known error rates implicate admissibility under both Federal Rule of Evidence 702 and Federal Rule 703. (Rule 702 is the locus of the scientific validity standard adopted in Daubert v. Merrell Dow Pharmaceutics, and Rule 703 states the common law, ad hoc balancing test for virtually all evidence.) I do not think the opinion is as clear as it could have been on which error rate pertained to what. Nevertheless, it is encouraging that the court recognized that two parameters are necessary to describe the accuracy of a procedure that classifies items or people into two categories (liar or truth teller).

But Daubert's list of factors extends beyond error rates, and the Semrau court’s handling of the other Daubert subissues also merits a mixed review. First, the court suggested that fMRI lie detection satisfied Daubert’s criteria for testing and peer review. It referred to “several factors in Dr. Semrau's favor,” namely:
“[T]he underlying theories behind fMRI-based lie detection are capable of being tested, and at least in the laboratory setting, have been subjected to some level of testing. It also appears that the theories have been subjected to some peer review and publication.” Semrau, 2010 WL 6845092, at *10. The Government does not appear to challenge these findings, although it does point out that the bulk of the research supporting fMRI research has come from Dr. Laken himself.
The suggestion that these factors favor the defendant treats Daubert’s references to testing, peer review, and publication rather superficially. That a scientific theory is “capable of being tested” tells us almost nothing about the validity of the theory. The theory that in the year 2075, the moon will turn into a blob of green cheese is capable of being tested, but that does not help validate it today. Likewise, the mere existence of peer reviewed publications means nothing without examining the content of the publications and the reactions to them in the scientific literature.

The court came closer to addressing the true Daubert issue of whether peer reviewed publications,  considered as a whole, validate a technique or theory when it responded to defendant’s argument that the district court was overly concerned with the realism of validity studies.  In that context, the court of appeals quoted the caveat in one fMRI study that:
This study has several factors that must be considered for adequate interpretation of the results. Although this study attempted to approximate a scenario that was closer to a real-world situation than prior fMRI detection studies, it still did not equal the level of jeopardy that exists in real-world testing. The reality of a research setting involves balancing ethical concerns, the need to know accurately the participant's truth and deception, and producing realistic scenarios that have adequate jeopardy.... Future studies will need to be performed involving these populations.
But even this mention of the content of one study does not explain why the experiments are inadequate to demonstrate validity. Why would it be harder to detect a lie that has grave consequences to the subject of the laboratory experiment or field study than one that has more trivial consequences?

Second, the court of appeals wrote that the “controlling standards factor” had not been satisfied because “[w]hile it is unclear from the testimony what the error rates are or how valid they may be in the laboratory setting, there are no known error rates for fMRI-based lie detection outside the laboratory setting, i.e., in the ‘real-world’ or ‘real-life’ setting.” But what does the realism of laboratory experiments have to do with the existence of a clear protocol for gathering and interpreting data? Naturally, if a test is not standardized, it is hard to ascertain its error rate—a point that has been prominent in debates over fingerprinting. And, if the tester departs slightly from the standard test protocol, the probative value of the test should be questioned under Rule 403. But the issue of external validity should not be confused with the issue of whether standards are in place for administering a test.

Finally, the court implied that without realistic field testing, there could be no general scientific acceptance of a method of lie detection in the forensic setting. This may be true, but all empirical studies pertain to particular times, places, and subjects. Deciding what generalizations are reasonable or generally accepted depends on understanding the phenomena in question. Can laboratory experiments alone show that certain factors tend to affect the accuracy of eyewitness identifications? For years, many experimental psychologists seemed willing to accept forensic applications of laboratory results that lacked complete realism. The ability of fingerprint examiners to match true pairs of prints and exclude false pairs of prints can be demonstrated in laboratory studies with artificially created pairs. Applying the error rates from such laboratory experiments to actual forensic setting could well be hazardous, but the experiments still prove that there is information that analysts can use to make valid judgments. In that situation, it is doubtful that raising the stakes of a judgment will render the technique invalid.

The Semrau court does not pinpoint the source of its discomfort with pure laboratory experiments. As we have just seen, a court should not assume that laboratory experiments never can establish validity of a technique as applied to casework. However, in the case of fingerprint identification, it seems clear enough that the prints do not change depending on whether they are deposited in the course of a crime or produced at another location. The fMRI data might well be different when generated under fully realistic circumstances. As a result, proving that there is detectable brain activity specific to conscious deception under low stakes conditions might not establish that the same pattern arises under high stakes conditions. Without a generally accepted theory of underlying mechanisms to justify extrapolations to the usual conditions of casework, low stakes laboratory findings may not suffice show general acceptance of validity under those conditions.

In sum, Semrau should not be read as establishing that the existence of testability carries any significant weight in favor of admission, that publication in a peer reviewed journal necessarily demonstrates validity, that a lack of complete realism in laboratory studies proves that there  are no “controlling standards” in practice, or that only field studies can establish general acceptance.

These concerns about the wording of the opinion notwithstanding, the problem of generalizing from the laboratory studies to the conditions of the Semrau case are substantial, and the court’s conclusion is difficult to dispute.

Wednesday, 26 September 2012

True Lies: fMRI Evidence in United States v. Semrau

This month, the U.S. Court of Appeals for the Sixth Circuit issued an opinion on “a matter of first impression in any jurisdiction.” The case is United States v. Semrau, No. 11-5396, 2012 WL 3871357 (6th Cir. Sept. 7, 2012). Its subject is the admissibility of the latest twist, the ne plus ultra, in lie detection—functional magnetic resonance imaging (fMRI).

In several ways, the case resembles what may well be the single most cited case on scientific evidence—namely, Frye v. United States, 293 F. 1013 (D.C. Cir. 1923). Frye instituted a special test for admitting scientific evidence. In Frye, a defense lawyer asked a psychologist, Dr. William Moulton Marston, who had developed and published studies of a systolic blood pressure test for conscious deception, to examine a young man accused of murdering a prominent physician. Dr. Marston came to Washington and was prepared to testify that the accused was truthful in retracting his confession to the murder. The trial court would not hear of it. The jury convicted. The defendant appealed. In a short opinion pregnant with implications, the Court of Appeals for the District of Columbia affirmed the exclusion of the expert’s opinion that the defendant was not lying to him.

In United States v. Semrau, defense counsel invited Dr. Steven Laken to examine the owner and CEO of two firms accused of criminal fraud in billing Medicare and Medicaid for psychiatric services that the firm supplied in nursing homes. Like Marston, Dr. Laken, had invented and published on an impressive method of lie detection. Following three sessions with the defendant, Dr. Laken concluded that the accused “was generally truthful as to all of his answers collectively.” As in Frye, the district court excluded such testimony. As in Frye, a jury convicted. As in Frye, the defendant appealed. As in Frye, the court of appeals affirmed.

Dr. Marston held degrees from Harvard in law and in psychology. He worked hard to develop and popularize psychological theories (and he created the comic book character, Wonder Woman). Like Marston, Dr. Laken is highly creative, productive, and enterprising. Dr. Laken started his scientific career in genetics and cellular and molecular medicine. He achieved early fame for discovering a genetic marker and developing a screening test for an elevated risk of a form of colon cancer. For that accomplishment, MIT’s Technology Review recognized him as one of the most important 35 innovators under the age of 35 and noted that “Laken believes his methods could spot virtually any illness with a genetic component, from asthma to heart disease.” I do not know if that happened. After four years as Director of Business Development and Intellectual Asset Management at Exact Sciences, a “molecular diagnostics company focused on colorectal cancer,” Laken left genetic science to found Cephos, “the world-class leader in providing fMRI lie detection, and in bringing fMRI technology to commercialization.”1/

Despite these parallels, Laken is not Marston, and Semrau is not Frye. For one thing, in Frye, the trial judge excluded the evidence without an explanation. In Semrau, the trial judge had a magistrate conduct a two-day hearing. Two highly qualified experts called by the government challenged the validity of Dr. Laken’s theories, and the magistrate judge wrote a 43-page report recommending exclusion of the fMRI testimony from the trial.2/

Furthermore, in Frye, there was no previous body of law imposing a demanding standard on the proponents of scientific evidence—the Frye court created from whole cloth the influential “general acceptance” test.3/ In Semrau, the court began with the Federal Rules of Evidence, ornately embroidered with the Supreme Court's opinions in Daubert v. Merrell Dow Pharmaceuticals and two related cases and with innumerable lower court opinions applying the Daubert trilogy. This legal tapestry requires a showing of “reliability” rather than “general acceptance,” and it usually involves attending to four or five factors relating to scientific validity enumerated in Daubert.3/ I want to look briefly at a few of these in the context of Semrau.

* * *

Even though the only judges to address fMRI-based lie detection (those in Semrau) have deemed it inadmissible under both the Daubert standard (and under the Frye criterion of general acceptance), Cephos continues to advise potential clients that “[t]he minimum requirements for admissibility of scientific evidence under the U.S. Supreme Court ruling Daubert v. Merrell Dow Pharmaceuticals, are likely met.” One can only wonder whether its “legal advisors,” such as Dr. Henry Lee (see note 1), are comfortable with Cephos’s reasoning that
According to a PubMed search, using the keywords ”fMRI” or “functional magnetic resonance imaging” yields over 15,000 fMRI publications. Therefore, the technique from which the conclusions are drawn is undoubtedly generally accepted.
The reasoning is peculiar, or at least incomplete. The sphygmomanometer that Dr. Marston used also was “undoubtedly generally accepted.” This pressure meter was invented in 1881, improved in 1896, and modernized in 1901, when Harvey Cushing popularized the device in the medical community. However, the acknowledged ability to measure systolic blood pressure reliably and accurately does not validate the theory—which predated Marston—that blood pressure is a reliable and valid indicator of conscious deception. Likewise, the number of publications about fMRI in general—and even particular evidence that it is a wonderful instrument with which to measure blood oxygenation levels in parts of the brain—reveals very little about the validity of the theory that these levels are well correlated with conscious deception. To be sure, there is more research on this association than there was on the blood pressure theory in Frye, but the Semrau courts were not overly impressed with applicability of the experimentation to the examination conducted in the case before it.4/

* * *

In addition to directing attention to general acceptance, Daubert v. Merrell Dow Pharmaceuticals identifies “the known or potential rate of error in using a particular scientific technique” as a factor to consider in determining “evidentiary reliability.” The Daubert Court took this factor from circuit court cases involving polygraphy and “voiceprints.” Unfortunately, the ascertainment of meaningful error rates has long confused the courts,5/ and the statistics in Semrau are not presented as clearly as one might hope.

According to Cephos, “[p]eer review results support high accuracy,” but this short statement begs vital questions. Accuracy under what conditions? How “high” is it? Higher for diagnoses of conscious deception than for diagnoses of truthfulness, or vice versa? The court of appeals began its description Semrau’s evidence on this score as follows:
Based on these studies, as well as studies conducted by other researchers, Dr. Laken and his colleagues determined the regions of the brain most consistently activated by deception and claimed in several peer-reviewed articles that by analyzing a subject's brain activity, they were able to identify deception with a high level of accuracy. During direct examination at the Daubert hearing, Dr. Laken reported these studies found accuracy rates between eighty-six percent and ninety-seven percent. During cross-examination, however, Dr. Laken conceded that his 2009 “Mock Sabotage Crime” study produced an “unexpected” accuracy decrease to a rate of seventy-one percent. ...
But precisely what do these “accuracy rates” measure? By “identify deception,” does the court mean that 71%, 86%, and 97% are the proportions of subjects who were diagnosed as deceptive out of those whom the experimenters asked to lie? If we denote a diagnosis of deception as a “positive” finding (like testing positive for a disease), then such numbers are observed values for the sensitivity of the test. They indicate the probability that given a lie, the fMRI test will detect it—in symbols, P(diagnose liar | liar), where “|” means “given.” The corresponding conditional error probability is the false negative probability P(diagnose truthful | liar) = 1 – sensitivity. It is the probability of missing the act of lying when there is a lie.

So far so good. But it takes two probabilities to characterize the accuracy of a diagnostic test. The other conditional probability is known as specificity. Specificity is the probability of a negative result when the condition is not present. In symbols that apply here, the specificity is P(diagnose truthful | truthful). Its complement, 1 – specificity, is the false positive, or false alarm, probability, P(diagnose liar | truthful). That is, the false alarm probability is the probability of diagnosing the condition as present (the subject is lying) when it is absent (the subject actually is not lying). What might the specificity be? According to the court,
Dr. Laken testified that fMRI lie detection has “a huge false positive problem” in which people who are telling the truth are deemed to be lying around sixty to seventy percent of the time. One 2009 study was able to identify a “truth teller as a truth teller” just six percent of the time, meaning that about “nineteen out of twenty people that were telling the truth we would call liars.” . . .
Why was this not a problem for Dr. Laken in this case? Well, the fact that the technique has a high false positive error probability (that it classifies most truthful subjects as liars) does not mean that it also has a high false negative probability (that it classifies most lying subjects as truthful). Dr. Laken conceded that the false positive probability, P(diagnose liar | truthful), is large (around 0.65, from the paragraph quoted immediately above). Indeed the reference to 6% accuracy for classifying liars (the technique’s sensitivity to lying), corresponds to a false positive probability of 100% – 6% = 0.94. The average figure for this false alarm probability, according to Dr. Laken’s statements in the preceding quoted paragraph, is lower, but it is still a whopping 0.65. Nevertheless, if the phrase “accuracy rates” in the first quoted paragraph refers to specificity, then the estimates of specificity that he provided are respectable. The average of 0.71, 0.86, and 0.97 is 0.85.

What do these numbers prove? One answer is that they apply only under the conditions of the experiments and only to subjects of the type tested in these experiments. The opinions take this strict view of the data, pointing out that the experimental subjects were younger than Semrau and that they faced low penalties for lying. Indeed, the court explained that
Dr. Peter Imrey, a statistician, testified: “There are no quantifiable error rates that are usable in this context. The error rates [Dr. Laken] proposed are based on almost no data, and under circumstances [that] do not apply to the real world [or] to the examinations of Dr. Semrau.”
These remarks go largely to the Daubert question. If the experiments are of little value in estimating an error rate in populations that would be encountered in practice, then the validity of the technique is difficult to gauge, and Cephos’s assurance that this factor weighs in favor of admissibility is vacuous. If there is no way to estimate the conditional error probability for the examination of Semrau, then it is hard to conclude that the test has been validated for its use in the case.

* * *

Fair enough, but I want to go beyond this easy answer. Psychologists often are willing to generalize from laboratory conditions to the real world and from young subjects (usually psychology students) to members of the general public. So let us indulge, at least arguendo, the heroic assumption that the ballpark figures for the specificity and the false alarm probability apply to defendants asserting innocence in cases like Semrau. On this assumption, how useful is the test?

Judging from the experiments as described in the court of appeals opinion, if Semrau is truthful in denying any intent to defraud, there is roughly a 0.85 probability of detecting it, and if he lies, there is maybe a 0.65 probability of misdiagnosing him as truthful. So the evidence—the diagnosis of truthfulness—is not much more probable when he is truthful than when he is lying. As such, the fMRI diagnosis of truthfulness has little probative value. (The likelihood ratio is .85/.65 = 1.3.)

That a diagnosis of deception is almost as probable for truthful subjects as for mendacious ones bears mightily on the Rule 403 balancing of prejudice against probative value. The court held that this balancing justified exclusion of Dr. Laken’s testimony, largely for reasons that I won’t go into.6/ It referred to questions about “reliability” in general, but it did not use the error probabilities to shed a more focused light on the probative value of the evidence.

However, it seems from the opinion that Dr. Laken offered at least one probability to show that his diagnosis was correct. The court noted that
Dr. Imrey also stated that the false positive accuracy data reported by Dr. Laken does not “justify the claim that somebody giving a positive test result ... [h]as a six percent chance of being a true liar. That simply is mathematically, statistically and scientifically incorrect.”
It is hard to understand what the “six percent chance” for “somebody giving a positive test result” had to do with the negative diagnosis (not lying) for Semrau. A jury provided with negative fMRI evidence (“He was not lying”) must decide whether the result is a true negative or a false negative—not what might have happened had there been a positive diagnosis.

As for the 6% solution, it is impossible to know from the opinion how Dr. Laken arrived at such a number for the probability that a subject is lying given a positive diagnosis. The conditional probabilities from the experiments run in the opposite direction. They address the probability of evidence (a diagnosis) given an unknown state of the world (a liar or a truthful subject). If Dr, Laken really opined on the probability of the state of the world (a liar) given the fMRI signals, then he either was naively transposing a conditional probability—a no-no discussed many times in this blog—or he was using Bayes’ rule. In light of Dr. Imrey’s impeccable credentials as a biostatistician and his unqualified dismissal of the number as “mathematically, statistically and scientifically incorrect,” I would not bet on the latter explanation.

Notes

1. If the firm’s website is any indication, it is not an equivalent leader in good grammar. Apparently seeking the attention of wayward lawyers, it advertises that “[i]f you or your client professes their innocence, we may provide pro bono consulting.” The website also offers intriguing reasons to believe in the company’s prowess: it is “represented by one of the top ten intellectual property law firms”; it has “been asked to present to the ... Sandra Day O’Connor Federal Courthouse”; and its legal advisors include Dr. Henry C. Lee (whose website includes “recent sightings of Dr. Lee.”). In addition to its lie-detection work, Cephos offers DNA testing, so perhaps I should not say that Dr. Laken has withdrawn entirely from genetic science.

2. The court of appeals buttressed its approval of the report with the observation that “Professor Owen Jones, who observed the hearing” and is on the faculties of law and biology at Vanderbilt University, stated in an interview with Wired, that the report was “carefully done.”

3. For elaboration, see David H. Kaye, David E. Bernstein & Jennifer L. Mnookin, The New Wigmore: A Treatise on Evidence—Expert Evidence (2d ed. 2011) http://www.aspenpublishers.com/product.asp?catalog_name=Aspen&product_id=0735593531

4. For a short discussion of validity in this context, see Francis X. Shen & Owen D. Jones, Brain Scans as Evidence: Truths, Proofs, Lies, and Lessons, 62 Mercer L. Rev. 861 (2011),

5. See David H. Kaye, David E. Bernstein & Jennifer L. Mnookin, The New Wigmore: A Treatise on Evidence—Expert Evidence (2d ed. 2011).

6. The court appeals wrote that “the district court did not abuse its discretion in excluding the fMRI evidence pursuant to Rule 403 in light of (1) the questions surrounding the reliability of fMRI lie detection tests in general and as performed on Dr. Semrau, (2) the failure to give the prosecution an opportunity to participate in the testing, and (3) the test result's inability to corroborate Dr. Semrau's answers as to the particular offenses for which he was charged.”