\(~\)
The problems associated with conducting multiple hypothesis tests are most pronounced in fields like genomics, where it is common to have data on thousands of biomarkers that each might be associated with a particular outcome.
The data provided below come a study seeking to find biomarkers that can identify people with early-stage Parkinson’s when the disease is most treatable. The researchers collected blood samples from 50 individuals with early-stage Parkinson’s as well as 55 similarly aged controls who did not have the disease.
This data set is large, so it might take a minute or so for you to load.
parkinsons_biomarkers = read.csv("https://remiller1450.github.io/data/parkinsons_biomarkers.csv")
Next, we’ll perform a two-sample \(t\)-test comparing the mean value of the 15,000 biomarkers across the early-stage Parkinson’s patients and controls. To do this we’ll use a programming tool known as a for loop to repeat the same testing procedure using a different column from the data frame:
p_values = numeric(15000)
for(i in 1:15000){
p_values[i] = t.test(parkinsons_biomarkers[,i] ~ parkinsons_biomarkers$status)$p.value
}
The vector p_values contains the \(p\)-value from each of these 15,000
two-sample \(t\)-tests.
\(~\)
As you should expect, running 15,000 two-sample \(t\)-tests creates the potential for a large number of Type 1 errors (false positives). In our lecture we learned about two ways to address this:
Both of these methods are implemented via the p.adjust()
function, which is demonstrated below for a small vector of example
\(p\)-values:
## Five example p-values ranging from 0.001 to 0.2
example_p_values = c(0.001, 0.005, 0.02, 0.04, 0.2)
## Bonferroni Adjusted p-values
p.adjust(example_p_values, method = "bonferroni")
## [1] 0.005 0.025 0.100 0.200 1.000
## FDR-adjusted p-values
p.adjust(example_p_values, method = "fdr")
## [1] 0.00500000 0.01250000 0.03333333 0.05000000 0.20000000
In this scenario we can easily count the significant \(p\)-values manually, but for larger data
sets (like the Parkinson’s example) it is helpful to know how to rely on
R to do the counting:
## Count p-values less than 0.05 for Bonferroni
example_adjusted_pvs = p.adjust(example_p_values, method = "bonferroni")
sum(example_adjusted_pvs <= 0.05)
## [1] 2
Question #1: For this question you should use the
Parkinson’s biomarkers data, more specifically the \(p\)-values stored in p_values
from the lab’s introduction.
R to count how many \(p\)-values were less than 0.01 (see the
example immediately preceding this question). Comparing your results
with your answer to Part A, does this seem like evidence that at least
some biomarkers in the data are associated with Parkinson’s disease?
Briefly explain your reasoning.\(~\)
A study conducted using an advanced driving simulator had \(n=19\) participants perform a 45-minute simulated drive under six different experimental conditions involving various combinations of alcohol, cannabis, and placebo:
Experimental conditions were randomly assigned across six different study sessions and separated by washout periods of at least 10 days.
One research question is whether a participant’s lateral control (lane-keeping ability) significantly worsened under various types of intoxication. Driving researchers assess lateral control using standard deviation of lane position (SDLP). Higher values of SDLP reflect greater variability in a driver’s position within their lane, and thus worse lateral control (more side to side movement).
The table below shows each study participant’s SDLP during an interstate road segment from their placebo-placebo drive, along with additional columns indicating whether their SDLP increased under the other experimental conditions, with “1” representing an increase and “0” representing a decrease or equivalent value.
| Participant | Baseline SDLP | Placebo/Low THC | Placebo/High THC | Alcohol/Placebo | Alcohol/Low THC | Alcohol/High THC |
|---|---|---|---|---|---|---|
| 1 | 14.2 | 0 | 0 | 1 | 1 | 1 |
| 2 | 17.9 | 0 | 1 | 1 | 1 | 0 |
| 3 | 19.9 | 1 | 0 | 1 | 1 | 0 |
| 4 | 11.4 | 0 | 0 | 1 | 1 | 1 |
| 5 | 18.3 | 1 | 1 | 0 | 0 | 1 |
| 6 | 18.5 | 0 | 1 | 1 | 1 | 0 |
| 7 | 15.8 | 1 | 1 | 1 | 1 | 1 |
| 8 | 15.9 | 1 | 1 | 1 | 1 | 0 |
| 9 | 15.8 | 0 | 0 | 1 | 1 | 0 |
| 10 | 15 | 1 | 1 | 0 | 1 | 0 |
| 11 | 16 | 1 | 1 | 1 | 1 | 1 |
| 12 | 14.7 | 0 | 1 | 1 | 1 | 1 |
| 13 | 15.3 | 1 | 1 | 1 | 1 | 1 |
| 14 | 17.4 | 1 | 1 | 0 | 1 | 1 |
| 15 | 19.6 | 0 | 0 | 1 | 1 | 1 |
| 16 | 16.9 | 1 | 1 | 0 | 0 | 1 |
| 17 | 15.9 | 1 | 0 | 1 | 1 | 0 |
| 18 | 14.9 | 1 | 1 | 1 | 0 | 1 |
| 19 | 15.1 | 1 | 1 | 1 | 1 | 1 |
| Total | 12 | 13 | 15 | 16 | 12 |
Question #2: For this question you should use data from the table provided above.
R to perform a test of the hypothesis \(H_0: p =0.5\) using an exact binomial test
(noting the small sample size violates the large counts assumption of
the one-sample Z-test). Report the \(p\)-value of this test and a one-sentence
conclusion.c(1, 2, 3, 4, 5) replacing the numbers 1-5 with the \(p\)-values.