Showing posts with label P-value. Show all posts
Showing posts with label P-value. Show all posts

Wednesday, 11 July 2012

More on Statistical Reasoning and the Higgs Boson

A posting of July 6, "The Probability that the Higgs Boson Has Been Discovered," mentioned the transposition of a p-value in stories in the popular press about the discovery of what is likely to be the Higgs Boson. Professor Dennis Lindley, a major figure in the development of Bayesian methods (and known to some readers of this blog as the author of a classic paper on using them to identify glass fragments) posed a few questions on the experiment via the list server of the International Society for Bayesian Analysis. One highly informed set of answers came from Louis Lyons (organiser of PHYSTAT series of meetings, and a member of CMS Collaboration at CERN). The following is a slightly edited version of the comments of Lindley (DL) and Lyons (LL). The comments presuppose knowledge of the meaning of a p-value, a likelihood ratio, Bayes' rule, and the divide between frequentists and Bayesians. (The original text as well as many other interesting messages are at http://bayesian.org/forums/news/3648.)

DL:
Specifically, the news referred to a confidence interval with 5-sigma limits.

LL:
The test statistic we use for looking at p-values is basically the likelihood ratio for the two hypotheses (H_0 = Standard Model (S. M.) of Particle Physics, but no Higgs; H_1 = S.M with Higgs). A small p_0 (and a reasonable p_1) then implies that H_1 is a better description of the data than H_0. This of course does not prove that H_1 is correct, but maybe Nature corresponds to some H_2, which is more like H_1 than it is like H_0. Indeed in principle data will never prove a theory is true, but the more experimental tests it survives, the happier we are to use it -- e.g. Newtonian mechanics was fine for centuries till the arrival of Relativity.

In the case of the Higgs, it can decay to different sets of particles, and these rates are defined by the S.M.  We measure these ratios, but with large uncertainties with the present data. They are consistent with the S.M. predictions, but it could be much more convincing with more data. Hence the caution about saying we have discovered the Higgs of the S.M.

DL:
Five standard deviations, assuming normality, means a p-value of around 0.0000005. A number of questions spring to mind.

1.  Why such an extreme evidence requirement? We know from a Bayesian perspective that this only makes sense if (a) the existence of the Higgs boson (or some other particle sharing some of its properties) has extremely small prior probability and/or (b) the consequences of erroneously announcing its discovery are dire in the extreme. Neither seems to be the case, so why 5-sigma?

LL:
This is an unfortunate tradition, that is used more readily by journal editors than by Particle Physicists. Reasons are
a) Historically we have had 3 and 4 sigma effects that have gone away

b) The 'Look Elsewhere Effect' (LEE). We are worried about the chance of a statistical fluctuation mimicking our observation, not only at the given mass of 125 GeV but anywhere in the spectrum. The quoted p-values are 'local' i.e. the chance of a fluctuation at the observed mass. Unfortunately the LEE correction factor is not very precisely defined, because of ambiguities about what is meant by 'elsewhere'

c) The possibility of some systematic effect (characterised by a nuisance parameter) being more important than allowed for in the analysis, or even overlooked - see the recent experiment at CERN which claimed that neutrinos travelled faster than the speed of light.

d) A subconscious use of Bayes Theorem to turn p-values into probabilities about the hypotheses.
All the above vary from experiment to experiment, so we realise that it is a bit unfair to use the same standard for discovery for all analyses. We prefer just to quote the p-values (or whatever).

DL:
2. Rather than ad hoc justification of a p-value, it is of course better to do a proper Bayesian analysis.  Are the particle physics community completely wedded to frequentist analysis?

LL:
No we are not anti-Bayesian, and indeed our test statistics is a likelihood ratio. If you like, you can regard our p-values as an attempt to calibrate the meaning of a particular value of the likelihood ratio.

We actually recommend that for parameter determination at the LHC, it is useful to compare Bayesian and Frequentist methods. But for comparing hypotheses (e.g. an experimental distribution is fitted by H_0 = a smooth distribution; or by H_1 = a smooth distribution plus a localised peak), we are worried about what priors to use for the extra parameters that occur in the alternative hypothesis.We would welcome advice.

DL:
3. We know that given enough data it is nearly always possible for a significance test to reject the null hypothesis at arbitrarily low p-values, simply because the parameter will never be exactly equal to its null value. And apparently the LHC has accumulated a very large quantity of data. So could even this extreme p-value be illusory?

LL:
We are aware of this. But in fact, although the LHC has accumulated enormous amounts of data, the Higgs search is like looking for a needle in  a haystack. The final samples of events that are used to look for the Higgs contain only tens to thousands of events.

These and related issues are discussed to some extent in my article "Open statistical issues in Particle Physics", Ann. Appl. Stat. Volume 2, Number 3 (2008), 887-915. It is supposed to be statistician-friendly.

Friday, 6 July 2012

The Probability that the Higgs Boson Has Been Discovered

Surely everyone has heard of the probable discovery of the Higgs boson. But what does it have to do with  forensic science or law? It is a reminder that the "prosecutor's fallacy" is not limited to prosecutors or courtrooms. Reports in the popular press by skilled physicists and science writers trying to explain this impressive discovery are replete with a messy form of the transposition fallacy. Here is an example from an otherwise excellent report by physicist Lawrence Krauss in Slate magazine:
One can in fact quantify the likelihood that the observations are mistaken and that the events are actually background noise mimicking a real signal. Each experiment quotes a likelihood of very close to “5 sigma,” meaning the likelihood that the events were produced by chance is less than one in 3.5 million. Yet in spite of this, the only claim that has been made so far is that the new particle is real and “Higgs-like.”
Likewise, Nature announced "just a 0.00006% probability that the result is due to chance." The New York Times reported that "the likelihood that their signal was a result of a chance fluctuation was less than one chance in 3.5 million, 'five sigma,' which is the gold standard in physics for a discovery," attributing the statement to CERN's physicists.

How is this (mis)reporting related to the transposition fallacy? Well, sigma (σ) stands for standard deviation, and 5σ means 5 standard deviations from the value expected if the measurements were just noise. For a normal distribution, results this extreme or more extreme would be seen in pure noise a small fraction of the time. The tiny figures quoted above are estimates of that fraction. The fraction is the statistician's p-value, P(>5σ | noise), and it is on the order of 10-6. In plain English (and one bit of Greek), the probability of data of more than 5σ given that they are just noise is on the order of one in a million. So the observations would be very surprising if they were just noise.

But the probability that they actually are noise is an inverse probability, P(noise | data). That probability depends on the likelihoods P(5σ | noise) and P(5σ | signal) as well as on the prior probability, P(noise). The p-value itself does not generally "quantify the likelihood that the observations are mistaken and that the events are actually background noise mimicking a real signal." It does not specify the "probability that the result is due to chance." If one wants to quantify the probability that the data are a real signal rather than noise, then, for better or worse, one must turn to Bayes' rule.

References (for physicists)

- Giulio D’Agostini, Bayesian Reasoning in High Energy Physics, CERN Yellow Report 99-03, July 1999
- Giulio D'Agostini, Probability and Measurement Uncertainty in Physics: A Bayesian Primer (1995)

A couple of other blogs (and one newspaper) making the same point

- http://www.r-bloggers.com/the-higgs-boson-sigma-5-and-the-concept-of-p-values/
- http://understandinguncertainty.org/higgs-it-one-sided-or-two-sided
- http://randomastronomy.wordpress.com/2012/07/04/higgs-boson-discovery-and-how-to-not-interpret-p-values/
- http://blog.carlislerainey.com/2012/07/07/innumeracy-and-higgs-boson/
- http://understandinguncertainty.org/explaining-5-sigma-higgs-how-well-did-they-do#comment-1449

- http://online.wsj.com/article/SB10001424052702303962304577509213491189098.html

Postscript

Professor Dennis Lindley, a major figure in the development of Bayesian methods (and known to some readers of this blog as the author of a classic paper on using them to identify glass fragments) posed a few questions on the Higgs boson experiment via the list server of the International Society for Bayesian Analysis. One well informed set of answers came from Louis Lyons (organiser of PHYSTAT series of meetings, and a member of CMS Collaboration at CERN). I posted a slightly edited version on July 11 under the title "More on Statistical Reasoning and the Higgs Boson." The full text of these and various other interesting messages is at http://bayesian.org/forums/news/3648.

Friday, 26 August 2011

The Transposition Fallacy in Matrixx Initiatives, Inc. v. Siracusano: Part II

The previous posting promised a simple example that would demonstrate the fallacy in claims such this one:
For a p-value of .09, the odds of observing the AER [adverse event report] is 91 percent divided by 9 percent. Put differently, there are 10-to-1 odds that the adverse effect is “real” (or about a 1 in 10 chance that it is not).
Brief of Amici Curiae Statistics Experts Professors Deirdre N. McCloskey and Stephen T. Ziliak in Support of Respondents, Matrixx Initiatives, Inc. v. Siracusano, 131 S.Ct. 1309 (2011) (No. 09-1156).

Here is one such example. A bag contains 100 coins. One of them is a trick coin with tails on both sides; the other 99 are biased coins that have a 0.3 chance of coming up tails and a 0.7 chance of coming up heads. I pick one of these coins at random and flip it twice, obtaining two tails. On the basis of only this sample data (the two tails), you must decide which type of coin I picked. The p-value with respect to the “null hypothesis” (N) that the coin is a normal (albeit biased) heads-tails one is the probability of seeing two tails in the two tosses: p = 0.3 x 0.3 = 0.09. Should you reject the null hypothesis N and conclude that I flipped the unique tails-tails coin? Are the odds for this alternative hypothesis (A) 10:1, as the brief of the statistical experts asserts?

Of course not. Just consider repeating this game over and over. Ninety-nine percent of the time, you would expect me to pick a heads-tails coin. In 9% of those cases, you expect me to get tails-tails on the two tosses (9% x 99% = 8.91%). The other way to get tails-tails on the tosses is to pick the tails-tails coin. You expect this to happen about 1% of the time. Thus, the odds of the tails-tails coin given the data on the outcome of the tosses are 1% to 8.91% = 1:8.91, which is about 1:9. Despite the allegedly significant (in “practical, human, or economic” terms) p-value of 0.09, the alternative hypothesis remains improbable.

A more formal derivation uses Bayes' rule for computing posterior odds. Let tt be the event that the coin I picked produced the two tails when tossed (the data), and let "|" stand for "given that" or "conditioned on." Then Bayes' rule reveals that

Odds(A|tt) = L x Odds(A),

where L is the "likelihood ratio" given by P(tt|A) / P(tt|N) and Odds(A) are the odds prior to flipping the coin. The value of L is 1/.09 = 100/9. Hence,

Odds(A|tt) = (100/9) Odds(A).

Because there is only 1 trick coin and 99 normal coins in the bag, the prior odds of A are Odds(A) = 1:99. Hence, the posterior odds are Odds(A|tt) = (100/9)(1/99) = 100/891 = 1:8.91. In other words, the odds for the alternative hypothesis are only about 1:9 -- practically the opposite of the 10:1 odds quoted in the statistics experts' brief.

The lesson of this example is not that a statistic with a p-value of 0.9 always can be safely ignored. It is that the p-value, by itself, cannot be converted into a probability that the alternative hypothesis is true (“that the adverse effect is ‘real’”). Knowing that the two tails arise only 9% of the time when the head-tails coin is the cause does not imply that 9% is the probability that a heads-tails coin is the cause or that 91% is the probability that the tails-tails coin is the “real” cause. Statisticians have warned against this confusion of a p-value with a posterior probability time and again. The brief of "Amici Curiae Statistics Experts" thus brings to mind the old remark, "With friends like these, who needs enemies?" A more complete review of the brief is available at Nathan Schachtman's website (see Further Readings).

Further Reading

David H. Kaye et al., The New Wigmore, A Treatise on Evidence: Expert Evidence (2d ed. 2011).

Nathan A. Schachtman, The Matrixx Oversold, Apr. 4, 2011, http://schachtmanlaw.com/the-matrixx-oversold/

Friday, 19 August 2011

The Transposition Fallacy in Matrixx Initiatives, Inc. v. Siracusano: Part I

One might expect to hear phrases like “Not statistical significance there!” and “There is no way that anybody would tell you that these ten cases are statistically significant” hurled by a disgruntled professor at an underperforming statistics student. Yet, in January 2011, they came from the Supreme Court bench during the argument in Matrixx Initiatives, Inc. v. Siracusano.[1]

The issue before the Court was “[w]hether a plaintiff can state a claim under § 10(b) of the Securities Exchange Act and SEC Rule 10b-5 based on a pharmaceutical company's nondisclosure of adverse event reports even though the reports are not alleged to be statistically significant.” [2] In the case, the manufacturer of the Zicam nasal spray for colds issued reassuring press releases at a time when it was receiving case reports from physicians of loss of smell (anosmia) in Zicam users. The pharmaceutical company, Matrixx Initiatives, succeeded in getting a security fraud class action dismissed on the ground that the plaintiffs failed to plead “statistical significance.”

Because case reports are just a series of anecdotes, it is not immediately obvious how they could be statistically significant, but a determined statistician could compare the number of reports in the relevant time period to the number that would be expected under some model of the world in which Zicam is neither a cause nor a correlate of anosmia. If the observed number departed from the expected number by a large enough amount—one that would occur no more than about 5% of the time when the assumption of no association is true (along with all the other features of the model)—then the observed number would be statistically significant at the 0.05 level.

The Court rejected any rule that would require securities-fraud plaintiffs to engage in such statistical modeling or computation before filing a complaint. This result makes sense because a reasonable investor might want to know about case reports that do not cross the line for significance. Such anecdotal evidence could be an impetus for further research, FDA action, or product liability claims—any of which could affect the value of the stock. In rejecting a bright-line rule of p < 0.05, the Court made several peculiar statements about statistical significance and the design of studies, but these are not my subject for today. (An older posting, on March 25, has some comments on this issue.)

Instead, I want to look at a small part of an amicus brief from “statistics experts” filed on behalf of the plaintiffs. There is much in this brief, which really comes from two economists (or perhaps these eclectic scholars should be designated historians or philosophers of economics and statistics), with which I would agree (for whatever my agreement is worth). But I was shocked to find the following text in the “Brief of Amici Curiae Statistics Experts Professors Deirdre N. McCloskey and Stephen T. Ziliak in Support of Respondents”:
The 5 percent significance rule insists on 19 to 1 odds that the measured effect is real.26 There is, however, a practical need to keep wide latitude in the odds of uncovering a real effect, which would therefore eschew any bright-line standard of significance. Suppose that a p-value for a particular test comes in at 9 percent. Should this p-value be considered “insignificant” in practical, human, or economic terms? We respectfully answer, “No.” For a p-value of .09, the odds of observing the AER [adverse event report] is 91 percent divided by 9 percent. Put differently, there are 10-to-1 odds that the adverse effect is “real” (or about a 1 in 10 chance that it is not). Odds of 10-to-1 certainly deserve the attention of responsible parties if the effect in question is a terrible event. Sometimes odds as low as, say, 1.5-to-1 might be relevant.27 For example, in the case of the Space Shuttle Challenger disaster, the odds were thought to be extremely low that its O-rings would fail. Moreover, the Vioxx matter discussed above provides an additional example. There, the p-value in question was roughly 0.2,28 which equates to odds of 4 to 1 that the measured effect — that is, that Vioxx resulted in increased risk of heart-related adverse events — was real. The study in question rejected these odds as insignificant, a decision that was proven to be incorrect.

26. At a 5 percent p-value, the probability that the measured effect is “real” is 95 percent, whereas the probability that it is false is 5 percent. Therefore, 95 / 5 equals 19, meaning that the odds of finding a “real” effect are 19 to 1.

27. Odds of 1.5 to 1 correspond to a p-value of 0.4. That is, the odds of the measured effect being real would be 0.6 / 0.4, or 1.5 to 1.

28. Lisse et al., supra note 14, at 543-44.
Why is this explanation out of whack? The fundamental problem is that, within the framework of classical (Neyman-Pearson) hypothesis testing, hypotheses like “the adverse effect is real” or “a measured effect being real” do not have odds or probabilities attached to them. In Bayesian inference, statements like “the probability that the measured effect is ‘real’ is 95 percent, whereas the probability that it is false is 5 percent” are meaningful, but frequentist p-values play no role in that framework. Equating the p-value with the probability that a null hypothesis is true and regarding the complement of a p-value as the probability that the alternative hypothesis is true (that something is “real”) is known as the transposition fallacy. [2] That two “statistics experts” would rely on this crude reasoning to make an otherwise reasonable point is depressing.

The preceding paragraph is a little technical. Soon, I shall post a simple example that should make the point more concretely and with less jargon.

References

1. Transcript of Oral Argument, Matrixx Initiatives, Inc. v. Siracusano, 131 S.Ct. 1309 (2011) (No. 09-1156), 2011 WL 65028, at *12 & *16 (Kagan, J.).

2. Petition for Writ of Certiorari at i, Matrixx Initiatives, Inc. v. Siracusano, 131 S.Ct. 1309 (2011) (No. 09-1156), 2010 WL 1063936.

3. David H. Kaye, David E. Bernstein & Jennifer L. Mnookin, The New Wigmore: A Treatise on Evidence: Expert Evidence (2d ed. 2011).