\(~\)

Onboarding

The problems associated with conducting multiple hypothesis tests are most pronounced in fields like genomics, where it is common to have data on thousands of biomarkers that each might be associated with a particular outcome.

The data provided below come a study seeking to find biomarkers that can identify people with early-stage Parkinson’s when the disease is most treatable. The researchers collected blood samples from 50 individuals with early-stage Parkinson’s as well as 55 similarly aged controls who did not have the disease.

This data set is large, so it might take a minute or so for you to load.

parkinsons_biomarkers = read.csv("https://remiller1450.github.io/data/parkinsons_biomarkers.csv")

Next, we’ll perform a two-sample \(t\)-test comparing the mean value of the 15,000 biomarkers across the early-stage Parkinson’s patients and controls. To do this we’ll use a programming tool known as a for loop to repeat the same testing procedure using a different column from the data frame:

p_values = numeric(15000) 
for(i in 1:15000){        
  p_values[i] = t.test(parkinsons_biomarkers[,i] ~ parkinsons_biomarkers$status)$p.value  
}

The vector p_values contains the \(p\)-value from each of these 15,000 two-sample \(t\)-tests.

\(~\)

Lab

Adjusted \(p\)-values

As you should expect, running 15,000 two-sample \(t\)-tests creates the potential for a large number of Type 1 errors (false positives). In our lecture we learned about two ways to address this:

  1. Family-wise Type 1 error control via the Bonferroni adjustment, which limits probability of making one or more Type 1 errors across an entire collection of tests to \(\alpha\) (ie: a 5% chance of making one or more false discoveries)
  2. False discovery rate control, which limits the expected proportion of false positives to \(\alpha\) (ie: at most 5% of significant findings are expected to be false discoveries)

Both of these methods are implemented via the p.adjust() function, which is demonstrated below for a small vector of example \(p\)-values:

## Five example p-values ranging from 0.001 to 0.2
example_p_values = c(0.001, 0.005, 0.02, 0.04, 0.2)

## Bonferroni Adjusted p-values
p.adjust(example_p_values, method = "bonferroni")
## [1] 0.005 0.025 0.100 0.200 1.000
## FDR-adjusted p-values
p.adjust(example_p_values, method = "fdr")
## [1] 0.00500000 0.01250000 0.03333333 0.05000000 0.20000000
  • The Bonferroni adjustment is very conservative, allowing only two of five \(p\)-values to be statistically significant after adjustment due to the method trying to avoid any false positives.
  • FDR control is less conservative, allowing four of five \(p\)-values to be statistically significant due to the method allowing a proportion of significant findings to be false discoveries.

In this scenario we can easily count the significant \(p\)-values manually, but for larger data sets (like the Parkinson’s example) it is helpful to know how to rely on R to do the counting:

## Count p-values less than 0.05 for Bonferroni
example_adjusted_pvs = p.adjust(example_p_values, method = "bonferroni") 
sum(example_adjusted_pvs <= 0.05)
## [1] 2

Question #1: For this question you should use the Parkinson’s biomarkers data, more specifically the \(p\)-values stored in p_values from the lab’s introduction.

  • Part A: Recall that there were 15,000 biomarkers analyzed using two-sample \(t\)-tests. If the null hypothesis of no difference between groups were true for all of these tests, how many biomarkers would you expect to be significant at the \(\alpha = 0.01\) level without any adjustment for multiple testing?
  • Part B: Use R to count how many \(p\)-values were less than 0.01 (see the example immediately preceding this question). Comparing your results with your answer to Part A, does this seem like evidence that at least some biomarkers in the data are associated with Parkinson’s disease? Briefly explain your reasoning.
  • Part C: Use the Bonferroni adjustment to control the family-wise Type 1 error rate at 20% and report the number of biomarkers that are identified as statistically significant.
  • Part D: In your own words, what does it mean to control the family-wise Type 1 error rate at 20%? How often would expect a false positive finding when using such a procedure?
  • Part E: Now control the false discovery rate at 20% and report the number of biomarkers that are identified as statistically significant.
  • Part F: In your own words, what does it mean to control the false discovery rate at 20%? How often would expect a false positive finding when using such a procedure?
  • Part G: In the context of this application, what is a Type 2 error? Briefly explain what such an error entails in your own words.
  • Part H: Which of the two adjustment procedures considered in this question (Bonferroni vs. FDR control) is likely to produce more Type 2 errors? Briefly explain your reasoning.
  • Part I: Suppose a geneticist consults you to analyze these data and you have the option of using the Bonferroni adjustment, a false discovery rate control approach, or making no adjustment to any of the \(p\)-values when reporting your results. Which of these options do you consider most appropriate for this application? Briefly explain your reasoning.

\(~\)

Application

A study conducted using an advanced driving simulator had \(n=19\) participants perform a 45-minute simulated drive under six different experimental conditions involving various combinations of alcohol, cannabis, and placebo:

  1. Placebo - Placebo
  2. Placebo - Low THC cannabis
  3. Placebo - High THC cannabis
  4. Alcohol - Placebo
  5. Alcohol - Low THC cannabis
  6. Alcohol - High THC cannabis

Experimental conditions were randomly assigned across six different study sessions and separated by washout periods of at least 10 days.

One research question is whether a participant’s lateral control (lane-keeping ability) significantly worsened under various types of intoxication. Driving researchers assess lateral control using standard deviation of lane position (SDLP). Higher values of SDLP reflect greater variability in a driver’s position within their lane, and thus worse lateral control (more side to side movement).

The table below shows each study participant’s SDLP during an interstate road segment from their placebo-placebo drive, along with additional columns indicating whether their SDLP increased under the other experimental conditions, with “1” representing an increase and “0” representing a decrease or equivalent value.

Participant Baseline SDLP Placebo/Low THC Placebo/High THC Alcohol/Placebo Alcohol/Low THC Alcohol/High THC
1 14.2 0 0 1 1 1
2 17.9 0 1 1 1 0
3 19.9 1 0 1 1 0
4 11.4 0 0 1 1 1
5 18.3 1 1 0 0 1
6 18.5 0 1 1 1 0
7 15.8 1 1 1 1 1
8 15.9 1 1 1 1 0
9 15.8 0 0 1 1 0
10 15 1 1 0 1 0
11 16 1 1 1 1 1
12 14.7 0 1 1 1 1
13 15.3 1 1 1 1 1
14 17.4 1 1 0 1 1
15 19.6 0 0 1 1 1
16 16.9 1 1 0 0 1
17 15.9 1 0 1 1 0
18 14.9 1 1 1 0 1
19 15.1 1 1 1 1 1
Total 12 13 15 16 12

Question #2: For this question you should use data from the table provided above.

  • Part A: Consider the placebo/low THC dosing condition. If this combination had no influence on SDLP, you’d expect it to be equally likely for an individual to have a higher or lower SDLP on their dosed drive relative to their baseline drive. With this in mind, use R to perform a test of the hypothesis \(H_0: p =0.5\) using an exact binomial test (noting the small sample size violates the large counts assumption of the one-sample Z-test). Report the \(p\)-value of this test and a one-sentence conclusion.
  • Part B: Which type of decision error (Type 1 or Type 2) could have been made in the analysis you performed in Part A? How do you know the other type of error did not occur? Briefly explain.
  • Part C: Repeat the test in Part A to obtain a \(p\)-value for each of the five dosing combinations. Store these \(p\)-values in a vector. You can assemble it yourself using something like c(1, 2, 3, 4, 5) replacing the numbers 1-5 with the \(p\)-values.
  • Part D: Suppose the null hypothesis is correct for each of the five tests involved in Part C and that these tests are all independent. What is the probability of making at least one Type 1 error if a significance threshold of \(\alpha = 0.05\) is used? Hint: Use the formula from Slide 10 of our lecture.
  • Part E: Use the Bonferroni correction to achieve a family-wise Type 1 error rate of 5% across the entire set of hypothesis tests you’ve performed. What do you conclude after applying this revised significance threshold? Note you do not need to repeat any hypothesis tests, instead you can simply report your adjusted \(p\)-values and an assessment of which dosing conditions (if any) differed from baseline.
  • Part F: In your own words, briefly describe the benefits and drawbacks of using the Bonferroni adjustment in this application. That is, why might researchers want to use the Bonferroni adjustment? Why might they want to report unadjusted \(p\)-values?