Sunday, 16 September 2012

ENCODE’S “Functional Elements” and the CODIS Loci (Part I)

Last week I noted some of the hyperbolic headlines accompanying the coordinated publication of a large number of datasets from the ENCODE Project . The abstract of the top-level paper begins as follows:
The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions.1/
Hoping to decipher these sentences, I have been reading about gene regulation. This modest effort stems from more than academic curiosity. If the popular and even some of the scientific press is to be believed, ENCODE has exorcized “junk DNA” from the body of scientific knowledge.2/ The bright light suddenly shining on the “dark matter” of the genome (to introduce another sloppy metaphor)3/ raises a giant question mark for the criminal justice system. Law enforcement authorities have always insisted that the snippets of DNA used to generate DNA identification profiles are just nonfunctional "junk."4/ Now, according to New York Times science correspondent Gina Kolata,
As scientists delved into the “junk” — parts of the DNA that are not actual genes containing instructions for proteins — they discovered a complex system that controls genes. At least 80 percent of this DNA is active and needed. … [¶] … The thought before the start of the project, said Thomas Gingeras, an Encode researcher from Cold Spring Harbor Laboratory, was that only 5 to 10 percent of the DNA in a human being was actually being used.5/
This juxtaposition of percentages suggests that the scientific community has shifted from the view that “only 5 to 10 percent” of the genome is functional (“needed” for the organism to function normally) to a sudden realization that 80% falls into this category.

But the more I read, the clearer it became that this description of a sudden phase transition in science is wildly inaccurate. Johns Hopkins biostatistian Steve Salzberg, in a penetrating and provocative Simply Statistics podcast interview, describes the 80% figure touted in the ENCODE paper as irresponsible.6/ University of Toronto biochemist Lawrence Moran saw it a repeat of a similar, problematic performance five years ago, at the conclusion of the pilot phase of ENCODE.7/ Responding to criticism, ENCODE Project leader Ewan Birney explained the new knowledge this way:
After all, 60% of the genome with the new detailed manually reviewed (GenCode) annotation is either exonic or intronic, and a number of our assays (such as PolyA- RNA, and H3K36me3/H3K79me2) are expected to mark all active transcription. So seeing an additional 20% over this expected 60% is not so surprising.8/
“Not so surprising”? A whopping 60%—not a minor 5 or 10%—was already estimated to be “active”? What is going on here?

The answer lies in the definition of some key terms (like exons, introns, and transcription) and requires a rudimentary understanding of the fundamentals of gene expression and its regulation in human beings. This posting presents the essential terminology and concepts. Part II will apply them to explain what ENCODE’s “assign[ing] biochemical functions for 80% of the genome” means. Anyone who knows what RNA transcripts and transcription factors do can skip this first part (or can read it to let me know of my inaccuracies).

To avoid suspense, I shall lay out my conclusions here and now: (1) if ENCODE gives a clear number for a percentage of the genome that regulates genes—the promoters, enhancers, silencers, ncRNA “genes,” and so on—I have yet to find it; (2) this number is almost surely less than the 80% figure reported for functionality; and (3) “functional element” as defined by the ENCODE Project is not a term that has clear or direct implications for claims of the law enforcement community that the loci used in forensic identification are not coding and therefore not informative. Those claims of zero information are somewhat exaggerated, but that is another story. For now, I merely describe some basics of gene expression and regulation.

Genes make proteins. But how? There are three big steps (with many activities within each step): transcription; post-transcription modification and transportation; and translation. All involve RNA, a single-stranded molecule related to DNA, and proteins. The basic picture is
  • Transcription to precursor messenger RNA: DNA + proteins --> pre-mRNA (in nucleus)
  • Post-transcriptional modification and transportation: pre-mRNA + proteins and RNAs -> mature m-RNA (in cytoplasm)
  • Translation to protein: mRNA + tRNA and proteins --> expressed protein (in cytoplasm)
In the first big step, the base pairs of the gene are transcribed jot-for-jot into an RNA molecule (precursor messenger RNA, or pre-mRNA). In the second major step, the transcript is modified at its ends, edited to remove parts that do not code for the protein that will be made (splicing), and the mature messenger RNA (m-RNA) is moved outside the nucleus. In the third phase, another type of RNA (transfer RNA, or tRNA) stitches together individual amino acids in the order dictated by the m-RNA transcript to form a protein, thereby translating the DNA sequence mirrored in the mRNA into the amino-acid order of the protein. Translation occurs on a kind of microscopic workbench (a ribosome) made of yet another RNA (ribosomal RNA, or rRNA).

For all this to happen, the DNA, which lies tightly coiled in the chromosomes (in a protein-DNA matrix known as chromatin), must open up for transcription to occur. Thus, changes in the chromatin regulate transcription, and these changes can be brought about in a number of ways. Transcription factors (specialized proteins) bind to the DNA. The bound transcription factors then recruit an enzyme (RNA polymerase) that produces RNA. This occurs within a region of DNA, known as a promoter, near the start of the protein-coding DNA (the structural gene). The level of transcription is influenced by activator or repressor proteins that bind to still other small regions (enhancers and silencers, respectively) that also lie outside the structural gene. In short, chemical interactions that open or close the chromatin that houses the DNA and transcription factors regulate the first step in the DNA-to-protein process.

In the past decade, other mechanisms of regulation or control of gene expression have been discovered. Many DNA sequences are not transcribed into messenger RNA, but they are transcribed into a variety of other RNAs. These non-protein-coding DNA sequences can be thought of as genes for RNA. Courting confusion, they usually are called “noncoding” (ncDNA)—because they do not code for protein—but they certainly code for RNAs that are crucial to translation—rRNA and tRNA—and for other RNAs that affect transcription, translation, and DNA replication. So it turns out that the genome is abuzz with transcription-to-RNA activity and other events that feed into the expression of the (protein-)coding DNA.

Yet, this hardly means that every biochemical event along the DNA is functionally important. Some, perhaps many, non-mRNA transcripts are just “noise.” They may float around for a while, but they may not do anything except wither away. In addition, large segments of the DNA transcribed in the course of making mRNA appear in the initial transcript (the pre-mRNA) but never make it into mature mRNA. These unused parts of the pre-mRNA transcripts correspond to long stretches of DNA, known as introns, that interrupt the smaller coding parts—the exons—that are translated into proteins. The initially transcribed intronic parts are removed from the pre-mRNA in a process called RNA splicing. Most of the RNA from introns probably just dissipates.9/

All these terms are a mouthful, but armed with this basic understanding of genes, RNA, and proteins, we can see why the 80% figure does not mean what one might think. We shall also see that the estimated proportion of the genome that encodes the structure of proteins or regulates gene expression has not jumped from 5 or 10% to 80%.

Notes

1. Ian Dunham et al., An Integrated Encyclopedia of DNA Elements in the Human Genome, 489 Nature 57 (2012).

2. E.g., Elizabeth Pennisi, ENCODE Project Writes Eulogy for Junk DNA, 337 Science 1159 (2012).

3. E.g., Gina Kolata, Bits of Mystery DNA, Far From ‘Junk,’ Play Crucial Role, N.Y. Times, Sept. 5, 2012. In one respect, the "dark matter" metaphor misrepresents dark matter. The presence of dark matter is inferred from its gravitational effects on visible matter. The presence of noncoding DNA is known from experiments that detect and characterize it just as they do coding DNA. Perhaps the metaphor means that the sequence of “dark matter” DNA cannot be deduced from the structure of a protein made in a cell. This, however, is like saying that dark matter is matter than cannot be seen with the naked eye. And that is not what astronomers mean by dark matter.

4. E.g., House Committee on the Judiciary, Report on the DNA Analysis Backlog Elimination Act of 2000, 106th Cong., 2d Sess., H.R. Rep. No. 106-900(1), at 27 (“the genetic markers used for forensic DNA testing … show only the configuration of DNA at selected ‘junk sites’ which do not control or influence the expression of any trait.”); New York State Law Enforcement Council, Legislative Priorities 2012: DNA at Arrest, at 5, http://nyslec.org/pdfs/2012/1_DNA_2012.pdf (“The pieces of DNA that are analyzed for the databank were specifically chosen because they are ‘junk DNA.’).

5. Kolata, supra note 2.

6. Interview by Roger Peng with Steven Salzberg, podcast on Simply Statistics, Sept. 7, 2012, http://simplystatistics.org/post/31056769228/interview-with-steven-salzberg-about-the-encode (“Why do they feel a need to say that 80% of the genome is functional? … They know it’s not true. They shouldn’t say it. … You don’t distort the science to get into the headlines.”).

7. Laurence A. Moran, The ENCODE Data Dump and the Responsibility of Scientists, Sept. 6, 2012, http://sandwalk.blogspot.com/2012/09/the-encode-data-dump-and-responsibility_6.html (“This is, unfortunately, another case of a scientist acting irresponsibly by distorting the importance and the significance of the data.”).

8. Ewan Birney, ENCODE: My Own Thoughts, Sept. 5, 2011

9. Post-splicing processing of a small fraction of the RNA from introns can produce noncoding RNAs that may regulate protein expression. L. Fedorova1 & A. Fedorov, Puzzles of the Human Genome: Why Do We Need Our Introns?, 6 Current Genomics 589, 592 (2005).

I am grateful to Eileen Kane for explaining some of the molecular biology to me.

Friday, 7 September 2012

Trashing Junk DNA

You have seen the headlines:
  • Bits of Mystery DNA, Far From 'Junk,' Play Crucial Role (New York Times)
  • 'Junk DNA' Concept Debunked by New Analysis of Human Genome (Washington Post)
  • 'Junk DNA' Debunked (Wall Street Journal)
  • Breakthrough Study Overturns Theory of 'Junk DNA' in Genome (Guardian)
Or maybe you heard MSNBC report that the data from ENCODE "shows us living beyond our genes" --whatever that means -- or listened to CBC intone that "'Junk DNA has a purpose" -- sounds divine -- or saw the Independent's mishugina announcement that "Scientists Debunk 'Junk DNA' Theory to Reveal Vast Majority of Human Genes Perform a Vital Function!" -- like we did not know that genes were functional and important?

The level of hype here is phenomenal. (Some useful clarification can be found at the Nature News blog.) In the next few days, I hope to post some quick thoughts on what the ENCODE figures (like 80%) being bandied about for the "functional" or "biologically active" fraction of the human genome mean for the loci used in forensic DNA identification.


(If any readers have insights to share, post a comment or send me an email at kaye at alum.mit.edu, and I'll try to use them. I am still educating myself about some of the details of gene regulation and can use any help I can get.)

Tuesday, 31 July 2012

Supreme Court to Review DNA Swabbing on Arrest??

According to the SCOTUS blog,
Chief Justice John G. Roberts, Jr., calling tests of the DNA of individuals arrested by police 'a valuable tool for investigating unsolved crimes,' on Monday cleared the way for the state of Maryland to continue that practice until the Supreme Court can act on a challenge to its constitutionality. The Chief Justice’s four-page opinion is here. A Maryland state court ruling against the practice will remain on hold until the Justices take final action.
One should not read these words as stating that the stay is in effect until the Justices decide whether Maryland constitutionally can take DNA from mere arrestees. That would require two further actions by the Court—"granting cert" and extending the stay while the Court decides the case—both unusual events. The Court receives over 8,000 petitions per year asking it to issue writs of certiorari—orders for lower courts to send the record to the Supreme Court for its review. The court grants on the order of 100 of them. It takes only four votes to grant a petition. (It used to require five.) Justice Scalia once called wading through piles of petitions and supporting materials "the most ... onerous and ... uninteresting part of the job." [1]

Thus far, the Chief Justice has issued a order (on his authority as a Circuit Justice) temporarily blocking ("staying") the judgment of the Maryland Court of Appeals. The Court of Appeals judgment did not order the state to do anything (although its import hardly could be ignored). It reversed the decision of the state's intermediate appellate court (that had upheld the constitutionality of Maryland's DNA-on-arrest law) and remanded the case to that lower court for further proceedings. (I described some notable features of the original Maryland Court of Appeals opinion on April 26. [2])

The Chief Justice's order remains in effect only until the other Justices of the Supreme Court get around to voting on Maryland's petition for a writ of certiorari. At that point, one of three things will happen: either (1) the Justices will grant the petition and decide to continue the freeze on the Maryland judgment while the Court reviews the case; (2) the Justices will grant the petition and let the stay elapse while they hear the case; or (3) they will deny the petition and leave the judgment of Maryland's highest court undisturbed. [3]

Thus, the Court's "final action" might be merely to decide not to act on the merits of the challenge to the constitutionality of the Maryland law. Denying cert has no precedential value. But the Chief Justice's July 30 opinion predicts that the Court actually will review the case and issue an opinion that will uphold the constitutionality of the law. Because of the contentiousness of the constitutional question, the brief opinion is worth dissecting.

The Chief Justice begins with the observation that "there is a reasonable probability this Court will grant certiorari." He ought to know, but the reason he gives is not entirely convincing. He writes that:
Maryland’s decision conflicts with decisions of the U. S. Courts of Appeals for the Third and Ninth Circuits as well as the Virginia Supreme Court, which have upheld statutes similar to Maryland’s DNA Collection Act. ... The split implicates an important feature of day-to-day law enforcement practice in approximately half the States and the Federal Government. ... Indeed, the decision below has direct effects beyond Maryland: Because the DNA samples Maryland collects may otherwise be eligible for the FBI’s national DNA database, the decision renders the database less effective for other States and the Federal Government.
But this "split" is nothing like a split in the federal circuits on the constitutionality of the federal database law. That kind of split would throw a real monkey wrench into the operation of NDIS, the FBI's National DNA Index System. The split here only affects timing and a fraction of all DNA profiles. That is, for those individuals who are convicted anyway, not taking DNA on arrest in Maryland only delays the time at which their profiles go into the database. Once the offender profiles are entered, a weekly database trawl should link them to any profiles in the database of crime-scene samples. Of course, this delay is not without costs. For example, some arrestees will commit other crimes, up to and including murder, in the period between arrest and conviction.

With respect to arrestees who never are convicted of offenses that trigger inclusion in the database, the state loses the opportunity to trawl the crime-scene database for their DNA profiles. Some of these individuals might be connected to these unsolved crimes, but many will not be. Thus, the split does not shut down the database system. It does reduce its efficiency by an amount that is not clearly known. As the Chief Justice puts it, "the decision renders the database less effective."

Chief Justice Roberts also writes that "the decision below subjects Maryland to ongoing irreparable harm" because "[A]ny time a State is enjoined by a court from effectuating statutes enacted by representatives of its people, it suffers a form of irreparable injury." The latter quotation comes from the previous Chief Justice, who expressed this claim in New Motor Vehicle Bd. of Cal. v. Orrin W. Fox Co., 434 U. S. 1345, 1351 (1977) (REHNQUIST, J., in chambers). But the notion that every court order that blocks enforcement of a duly enacted law works an irreparable injury seems extravagant. Does the public suffer irreparable harm when someone on a Fort Lauderdale beach plays frisbee, flies a kite, attaches a hammock to a tree, or swims in long pants—all prohibited?

The more meaningful argument is that the Maryland ruling constitutes "an ongoing and concrete harm to Maryland’s law enforcement and public safety interests." The Chief Justice explains: "According to Maryland, from 2009—the year Maryland began collecting samples from arrestees—to 2011, 'matches from arrestee swabs [from Maryland] have resulted in 58 criminal prosecutions.'" But this statistic is wide of the mark. How many of these 58 prosecutions would the state have foregone had it been unable to enter the profiles at the point of the arrest rather than waiting until a conviction ensued?

In short, the Chief Justice is correct in stating that "in the absence of a stay, Maryland would be disabled from employing a valuable law enforcement tool for several months," but his opinion leaves unresolved the question of just how valuable it really is. This is a matter that surely will receive more attention if and when the full Court actually hears the case.

References

Wednesday, 25 July 2012

CODIS Loci Not Ready for Disease Prediction After All?

Last month, I noted the findings of a superior court in Vermont that "some of the CODIS loci have associations with identifiable serious medical conditions," making the scientific evidence "sufficient to overcome the previously held belief[s]" about the innocuous nature of the CODIS loci [1]. The judge based her conclusion in State v. Abernathy [2] that the CODIS loci now permit "probabilistic predictions of disease" on the unpublished views of biologist Greg Wray, who oversees the Center for Evolutionary Genomics and the DNA Sequencing Core Facility, within Duke University’s Institute for Genome Sciences and Policy.

A technical report accepted for publication in the Journal of Forensic Sciences seems to dispute these claims. Sara Katsanis, a staff researcher at the same Institute for Genome Sciences and Policy, and Jennifer Wagner, a research associate at the University of Pennsylvania’s Center for the Integration of Genetic Healthcare Technologies, searched the biomedical literature and genomic databases not only for associations with phenotypes in the current 13 loci used in offender databases, but also in ones that soon may be added to the system. They came up with “no evidence” that any particular CODIS single-locus genotypes “are indicative of phenotype.”

References

1. CODIS Loci Ready for Disease Prediction, Vermont Court Says, June 15, 2012.
2. State v. Abernathy, No. 3599-9-11 (Vt. Super. Ct. June 1, 2012).

Saturday, 14 July 2012

Going South with Shoeprint Testimony

On July 9, 10, and 12 ("If the Shoe Fits, You Must Not Calculate It"), I discussed the much maligned opinion in R. v. T., [2010] EWCA Crim. 2439. Professor Mike Redmayne kindly called to my attention another Court of Appeal opinion on shoe print evidence: R. v. South, [2011] EWCA Crim. 754. It indicates that an expert  who does not use the adjective "scientific" can present the "verbal equivalent" of a likelihood ratio when the precise value of the ratio is uncertain. Indeed, the witness can give a source probability as long as it is a personal judgment derived solely from individual experience rather than from systematically collected data on shoe prints. This situation is not the best of all possible worlds.

Students in Bournemouth found that their house had been burgled in the afternoon as two of them slept. “On the floor below the letterbox of the front door there were some envelopes which had footmarks on them. These envelopes were subsequently given to the police and they were forensically examined. The evidence concerning those footprints was adduced at the trial.” This was by no means the only evidence the police developed against Sergio South, who was known to them as a burglar, but it became fodder for his appeal.

The evidence from an FSS examiner (a Mr. Jones) resembled that of the FSS's Mr. Ryder in R. v. T. Both experts, quite reasonably, relied on size, pattern, and wear. Here, "this footprint was in agreement with the size, pattern, detailed alignment and degree of wear with the trainer of the appellant that had been seized from him upon arrest. The zigzag bar pattern and the curved tramline were similar, and the trainers, which were size 9, were consistent with the footprint which was of size 9 or 8 but not size 10."

As in R. v. T., the expert must have consulted the FSS likelihood ratio table of “verbal equivalents.” Defense counsel “submitted that ... Mr Jones had said that the evidence relating to the footprint was ‘moderately strong support’ for the proposition that the appellant's shoe had made the imprint on the envelopes.” The phrase “moderately strong” is reserved in the FSS table for likelihood ratios between 100 and 1000—the next rung up the ladder from the “moderate support” for the ratio of 100 in R. v. T.

So what distinguishes the cases? Surely not that South's feet were smaller or the likelihood ratio larger. Is it that Mr. Jones was asked on cross-examination where the verbal equivalent came from?  Defense counsel advised the Court of Appeal that "Mr Jones had said that this expression reflected a statistical probability of the footprint having been made by the shoes of the appellant which was considerably more than a 50 per cent probability, because the linguistic phrases used, such as 'weak or limited support' or 'extremely strong support', were based on probability which was itself based on a logarithmic scale."

If this description of the cross-examination is correct, the expert's testimony was less defensible that that in R. v. T.  It appears that Mr. Jones missed the point of the FSS's efforts to train analysts in estimating likelihood ratios. The raison d'etre for using likelihoods to arrive at a standardized expression for the strength of the evidence is to get away from testimony about source probabilities. Statements such as "a statistical probability of the footprint having been made by the shoes of the appellant ... was considerably more than a 50 per cent probability" are strictly verboten. The expert following the strength-of-evidence approach must confine himself to commenting on the degree to which the evidence supports the competing claims about the source of the impressions. It is the role of the jury, and not the business of the expert, to consider the probability of those claims. In addition the expert (or the defense counsel) did not appreciate the fundamental difference between probabilities (of hypotheses about the origin of the marks) and likelihood ratios (which measure the support the evidence gives to those hypotheses). Whether this foggy cross-examination satisfied R. v. T.'s call for more transparency about the origin of an expert's description of the strength of the trace evidence is questionable.

Nevertheless, and even though Mr. Jones testified as "a scientist" who "had worked as in this area since 1982," the court concluded that his presentation "did not transgress in any way the guidelines set down by this court in R v T." The crucial fact for the court was that "Mr Jones' evidence was based on his experience." The South court described R. v. T. as stating “that if a footwear examiner expressed a view that went beyond saying that the footwear could or could not make the mark concerned, the report should make it clear that the view is subjective and based on experience of the examiner, so that words such as ‘scientific’ used in making evaluations should not in fact be used because they would, before a jury, give an impression of a degree of precision and objectivity which is not present given the current state of expertise.”

That Mr. Jones referred to "a statistical probability" did not seem to worry the court. Neither did the court perceive any problem with testimony "that he encountered the type of footwear seized from the appellant in only 2 per cent of cases that he dealt with as a forensic examiner of footwear and footprints" and "that burglars frequently used sports trainers." If "2 per cent" is a summary of 27 years of unrecorded personal experiences, it is hardly a rigorously ascertained "statistical probability," although it is a statistic and it yields a probability. And if Mr. Jones's understanding of the sartorial preferences of burglars informed his perception of "moderately strong" trace evidence yielding a posterior probability of "considerably more than 50 per cent," then he was exceeding the bounds of his expertise as a careful observer of similarities and differences and a keen analyst of the significance of these similarities and difference in impressions.

In sum, South indicates that the strictures of R. v. T. are easily avoided. But the courts and the forensic science profession do better. They can implement a system of reporting and testifying that conveys opinions or information in terms of the strength of the evidence rather than the probability of source hypotheses. There is considerable support for this approach among forensic service providers in Europe. The English courts lag behind, as do both the forensic science profession and the courts in the United States.

Thursday, 12 July 2012

If the Shoe Fits, You Must Not Calculate It (Part III)

In R v T (see posts of July 9 and 10), the Court of Appeal for England and Wales was distressed that a footwear analyst used a database on the characteristics of shoes to verify his holistic impression that the forensic science evidence—the correspondence between the defendant’s shoes and the marks at a murder scene—constituted "a moderate degree of scientific evidence to support the view that the [shoes] had made the footwear marks." The court had this to say (in part) about the resort to the database:
Mr Ryder used the internal database of the FSS to examine the frequency of pattern. This recorded the number of shoes received by the FSS (in contradistinction to the number distributed within the United Kingdom ... ). The FSS database comprised approximately 0.00006 per cent of all shoes sold in a year. ...

It is evident from the way in which Mr Ryder identified the figures to be used in the formula for pattern and size that none has any degree of precision. The figure for pattern could never be accurately known. For example, there were only distribution figures for the United Kingdom of shoes distributed by Nike; these left out of account the Footlocker shoes and counterfeits. ...

More importantly, the purchase and use [of] footwear is also subject to numerous other factors such as fashion, counterfeiting, distribution, local availability and the length of time footwear is kept. A particular shoe might be very common in one area because a retailer has bought a large number or because the price is discounted or because of fashion or choice by a group of people in that area. There is no way in which the effect of these factors has presently been statistically measured; it would appear extremely difficult to do so, but it is an issue that can no doubt be explored for the future.

It is important to appreciate that the data on footwear distribution and use is quite unlike DNA. A person’s DNA does not change and a solid statistical base has been developed which enable accurate figures to be produced. Indeed as was accepted by Mr Ryder, the data for footwear sole patterns is a small proportion of what is in use and changes rapidly. [I]t would for these reasons be dangerous to use a straight statistical model.

Use of the FSS’s own database could not have produced reliable figures as it had only 8,122 shoes whereas some 42 million are sold every year. [T]he likelihood ratio calculated by using figures for the population as a whole is completely different from that calculated using the figures used by Mr Ryder based on the FSS database. There is also the further difficulty, even if it could be used for this purpose, that the data are the property of the FSS and are not routinely available to all examiners. It is only available in a particular case to an examiner appointed to consider the report of an FSS examiner.
The court’s uneasiness with the database involves two considerations—sample size and relevance. Each merits discussion.

Sample Size

The court’s dismissal of the sample as a mere "0.00006 per cent of all shoes sold in a year" stems from the intuition that accurate estimation of a population proportion requires a sample that constitutes a large fraction of the population of interest. This perception is common but statistically naïve. If the sample is random and the population is very large compared to the sample, the statistical uncertainty in the estimate depends on the absolute size of the sample, not the sample size relative to the population size.

To see this, imagine a huge container of randomly packed with marbles of two colors (20% blue, 80% red). I mean a huge container—almost half of an entire football stadium is filled with marbles! They go halfway up the height of the bleachers. If we pick a large sample (say, 10,000 marbles), it is likely to represent all the marbles in the stadium pretty well. It would be surprising if the proportion of blue marbles in the sample of 10,000 were very different from 20%.

Now we dump in truckloads more of well mixed marbles in the same 20-80 proportion of colors until the stadium is overflowing with blue and red marbles. The population of marbles is now three times what it was before. Must we triple the sample size to keep pace with the larger population?

Absolutely not. That the second sample is an even tinier fraction of the population than the first one is irrelevant. The same sample (size 10,000) will work as well for the packed stadium as it did for the partly full stadium. For large samples from very much larger populations, the precision of the sample estimate—of marble colors, shoe sizes, or what have you—essentially depends on the absolute size of the sample (actually, the square root of the sample size). The effect of population size is negligible. [1]

Relevance (Fit)

Even if sample size is not the issue here, does the sample relate to the population of interest? This is a matter of relevance, or what the US Supreme Court called “fit” in the famous American case of Daubert v. Merrell Dow Pharmaceuticals. In many early DNA cases, courts thought they needed to know the frequencies of DNA profiles in a defendant's ethnic or racial group although that plainly was not the relevant population. In one New York case, for instance, the court worried that the defendant's profile might be common in the defendant's home town of Shushtar, Iran—even though the alleged sexual assault took place in a wealthy, suburban community in New York [2].

A possible disconnect between population of shoes represented in the FSS database and the population of "innocent shoes" that, ideally, should be sampled means that the analyst should not overstate the value of the FSS data. Moreover, in some cases the available data might be too far afield to be particularly helpful, but usually having some systematically collected statistical information is better than having none, and a factfinder can appreciate the limitations in those background data.

Several articles on R v T make these points. Mike Redmayne, Paul Roberts, Colin Aitken, and Graham Jackson perspicaciously note that sample size is not the serious issue here. Rather,
The pertinent question is: which database is likely to provide the best comparators, relative to the task in hand? Shoe choice is influenced by social and cultural factors. Middle-aged university lecturers presumably buy different trainers to teenage schoolboys. And the shoes making the marks most often found at crime scenes are not the general public's most popular purchases. Consequently, the FSS database—comprising shoes owned by those who, like T, have been suspected of committing offences—may well be a more appropriate source of comparison data than national sales figures [3, pp. 354-55].
They add:
While the court was concerned that there “are, at present, insufficient data for a more certain and objective basis for expert opinion on footwear marks”, it is rarely helpful to talk about “objective” data in forensic contexts. Choice of data always involves a degree of judgement about whether a particular dataset is fit for purpose. At the same time, reliance on data is ubiquitous and inescapable. When the medical expert witness testifies that “I have never encountered such a case in forty years of clinical practice”, he is utilising data but he calls them “experience”, relying on memory rather than any formal database open to probabilistic calculations. It would be just as foolish to maintain that memory and experience are never superior to quantified probabilities in criminal litigation, as it would be to insist that memory and experience are always preferable to and should invariably displace empirical data and quantified probabilities in the courtroom.
There is a long running debate in psychology over the relative merits of clinical versus statistical prediction, but who can deny that a bad statistical model can give less accurate results than insightful clinicians? Nonetheless, “objective” inferences based on publicly accessible data have something going for them even when they are merely comparable to gestalt judgments. They can avoid cognitive bias in highly subjective and complex decision-making. Statistical models for deciphering complicated DNA mixtures or for gauging similarities in fingerprints have this appeal.

Objective and Subjective Probabilities

Another incisive article on R v T, by Charles Berger, John Buckleton, Christophe Champod, Ian W. Evett, and Graham Jackson, pursues the objective-subjective dichotomy. Some of their remarks could be read as suggesting that objective probabilities are ultimately subjective:
[W]henever we are making an inference from a sample the data are always an incomplete representation of the full picture; furthermore, their relevance is a matter of judgement and the uncertainty that concerned the Court is an unavoidable feature of such inference. The probability that is quoted then will inevitably be a personal probability and the extent to which the data influence that probability will depend on expert judgement. This is not a process that can be governed strictly by mathematical reasoning but this does not make it any less “scientific”: scientists are called on to exercise personal judgement in all aspects of their several pursuits [4, p. 45 (emphasis added)].
That judgment is ubiquitous in scientific reasoning is a welcome antidote to idealized visions of science, but whether “the probability that is quoted then will inevitably be a personal probability” depends on what probability is quoted. Unlike the Bayesian, the frequentist does not quote a probability for the truth of an inference from the sample to a population. The frequentist, upon finding that the sample proportion is 0.2 (the figure in the FSS database), reports that 0.2 is a reasonable estimate of the proportion in the population from which the sample was drawn (assuming, of course, that the sample was drawn at random and is reasonably large). The frequentist relies strictly on the sample data to go from the sample data to the population parameter. Whether the FSS shoes were drawn at random and the nature of the population from which they were drawn are crucial to frequentists, but they are not matters that the frequentist expert can discuss in terms of probabilities.

A Bayesian, on the hand, does not use just the sample proportion to estimate the population proportion. The Bayesian attaches probabilities to all possible prior beliefs about the population proportion and modifies these in view of the sample data. If the Bayesian’s prior beliefs were concentrated at 0.5 (that is, the Bayesian strongly believed before looking at the FSS database that half the population of shoes had the pattern in question), then the Bayesian might report an estimated population proportion at a point closer to 0.5 than 0.2—perhaps 0.4.

This does not make the frequentist’s estimate into a personal probability, and it does not address the problems of model selection that affect the Bayesian and the frequentist alike. Both the frequentist and the Bayesian estimates (of 0.2 and 0.4, respectively) pertain to the population from which the FSS database was drawn, presumably at random.

Whether that population is the best one to consider is a further question. Both the Bayesian and the frequentist might agree the locale in which the crime was committed is distinctive in ways that would make the population proportion smaller than that of the population from which the FSS acquired its shoes. In that event, both could present their estimates as conservative (likely to favor the defendant). But the uncertainty has not converted the "objective" 0.2 figure into a personal probability. Frequentist probabilities are conditioned on a specific model. Whether a model is reasonable, they would readily concede, is vital—but it is not a judgment to which they can or will assign a probability.

Berger et al. also observe that “[f]urthermore, the probabilities that the scientist is directed to address are always founded (even with DNA) on personal judgement. This is not a bad thing, it is an inescapable feature of science ... [3, p. 49].” That the scientist makes judgments is indeed inescapable. But does that make the probabilities or parameters that the scientist estimates personal and subjective? The objectivist would hold that it only means that (1) the scientist’s estimates of these quantities are based on assumptions; (2) the plausibility of the assumptions are matters of judgment; and (3) it is a good thing to make these assumptions explicit so that other scientists and legal factfinders can consider whether they are sufficiently reasonable to make the objective probabilities helpful.

In short, Berger et al. are right—a court should not get carried away with the distinction between objective and subjective probabilities. I would merely add that neither should courts act as if there are no differences between them.

References

1. Hans Zeisel & David H. Kaye, Prove It with Figures: Empirical Methods in Law and Litigation (1997).

2. People v. Mohit, 153 Misc.2d 22, 579 N.Y.S.2d 990 (Westchester Co. Ct. 1992)
3. Mike Redmayne, Paul Roberts, Colin Aitken, Graham Jackson, Forensic science evidence in questions, Crim. L.R. 2011(5) 347–356

4. Charles E.H. Berger, John Buckleton , Christophe Champod, Ian W. Evett, Graham Jackson, Evidence Evaluation: A Response to the Court of Appeal Judgment in R v T, Science and Justice 2011 51: 43–49

Wednesday, 11 July 2012

More on Statistical Reasoning and the Higgs Boson

A posting of July 6, "The Probability that the Higgs Boson Has Been Discovered," mentioned the transposition of a p-value in stories in the popular press about the discovery of what is likely to be the Higgs Boson. Professor Dennis Lindley, a major figure in the development of Bayesian methods (and known to some readers of this blog as the author of a classic paper on using them to identify glass fragments) posed a few questions on the experiment via the list server of the International Society for Bayesian Analysis. One highly informed set of answers came from Louis Lyons (organiser of PHYSTAT series of meetings, and a member of CMS Collaboration at CERN). The following is a slightly edited version of the comments of Lindley (DL) and Lyons (LL). The comments presuppose knowledge of the meaning of a p-value, a likelihood ratio, Bayes' rule, and the divide between frequentists and Bayesians. (The original text as well as many other interesting messages are at http://bayesian.org/forums/news/3648.)

DL:
Specifically, the news referred to a confidence interval with 5-sigma limits.

LL:
The test statistic we use for looking at p-values is basically the likelihood ratio for the two hypotheses (H_0 = Standard Model (S. M.) of Particle Physics, but no Higgs; H_1 = S.M with Higgs). A small p_0 (and a reasonable p_1) then implies that H_1 is a better description of the data than H_0. This of course does not prove that H_1 is correct, but maybe Nature corresponds to some H_2, which is more like H_1 than it is like H_0. Indeed in principle data will never prove a theory is true, but the more experimental tests it survives, the happier we are to use it -- e.g. Newtonian mechanics was fine for centuries till the arrival of Relativity.

In the case of the Higgs, it can decay to different sets of particles, and these rates are defined by the S.M.  We measure these ratios, but with large uncertainties with the present data. They are consistent with the S.M. predictions, but it could be much more convincing with more data. Hence the caution about saying we have discovered the Higgs of the S.M.

DL:
Five standard deviations, assuming normality, means a p-value of around 0.0000005. A number of questions spring to mind.

1.  Why such an extreme evidence requirement? We know from a Bayesian perspective that this only makes sense if (a) the existence of the Higgs boson (or some other particle sharing some of its properties) has extremely small prior probability and/or (b) the consequences of erroneously announcing its discovery are dire in the extreme. Neither seems to be the case, so why 5-sigma?

LL:
This is an unfortunate tradition, that is used more readily by journal editors than by Particle Physicists. Reasons are
a) Historically we have had 3 and 4 sigma effects that have gone away

b) The 'Look Elsewhere Effect' (LEE). We are worried about the chance of a statistical fluctuation mimicking our observation, not only at the given mass of 125 GeV but anywhere in the spectrum. The quoted p-values are 'local' i.e. the chance of a fluctuation at the observed mass. Unfortunately the LEE correction factor is not very precisely defined, because of ambiguities about what is meant by 'elsewhere'

c) The possibility of some systematic effect (characterised by a nuisance parameter) being more important than allowed for in the analysis, or even overlooked - see the recent experiment at CERN which claimed that neutrinos travelled faster than the speed of light.

d) A subconscious use of Bayes Theorem to turn p-values into probabilities about the hypotheses.
All the above vary from experiment to experiment, so we realise that it is a bit unfair to use the same standard for discovery for all analyses. We prefer just to quote the p-values (or whatever).

DL:
2. Rather than ad hoc justification of a p-value, it is of course better to do a proper Bayesian analysis.  Are the particle physics community completely wedded to frequentist analysis?

LL:
No we are not anti-Bayesian, and indeed our test statistics is a likelihood ratio. If you like, you can regard our p-values as an attempt to calibrate the meaning of a particular value of the likelihood ratio.

We actually recommend that for parameter determination at the LHC, it is useful to compare Bayesian and Frequentist methods. But for comparing hypotheses (e.g. an experimental distribution is fitted by H_0 = a smooth distribution; or by H_1 = a smooth distribution plus a localised peak), we are worried about what priors to use for the extra parameters that occur in the alternative hypothesis.We would welcome advice.

DL:
3. We know that given enough data it is nearly always possible for a significance test to reject the null hypothesis at arbitrarily low p-values, simply because the parameter will never be exactly equal to its null value. And apparently the LHC has accumulated a very large quantity of data. So could even this extreme p-value be illusory?

LL:
We are aware of this. But in fact, although the LHC has accumulated enormous amounts of data, the Higgs search is like looking for a needle in  a haystack. The final samples of events that are used to look for the Higgs contain only tens to thousands of events.

These and related issues are discussed to some extent in my article "Open statistical issues in Particle Physics", Ann. Appl. Stat. Volume 2, Number 3 (2008), 887-915. It is supposed to be statistician-friendly.