R\(~\)
The main theme of this lab is answering the question “which one-sample test should I use?”
As a general decision framework, before running any statistical test we should consider the following:
| Outcome Type | Sample Characteristics | Recommended Test |
|---|---|---|
| Categorical | Large sample (at least 10 observations in each category) | One-sample Z-test - prop.test() |
| Categorical | Small sample (fewer than 10 observations in one category) | Exact Binomial Test - binom.test() |
| Quantitative | Population approximately Normal OR \(n \geq30\) | One-sample T-test - t.test() |
| Quantitative | Small sample (\(n<30\)) and strong evidence of non-Normality | Wilcoxon Signed-Rank Test -
wilcox.test() |
We have covered the one-sample Z-test and T-test in our previous two lectures, and this lab will briefly cover the exact binomial and Wilcoxon signed-rank tests.
After introducing the R implementation of each of these
tests using examples and practice questions, the lab features two
questions to practice applying a consistent statistical workflow:
\(~\)
Recall the one-sample \(Z\)-test involves two primary steps:
The prop.test() function in R can be used
to find the \(p\)-value produced by the
one-sample Z-test. The code below uses prop.test() to
replicate the one-sample \(Z\)-test on
the infant toy choice data from our previous lecture:
prop.test(x = 14, n = 16, p = 0.5, alternative = "greater", correct = FALSE)
##
## 1-sample proportions test without continuity correction
##
## data: 14 out of 16, null probability 0.5
## X-squared = 9, df = 1, p-value = 0.00135
## alternative hypothesis: true p is greater than 0.5
## 95 percent confidence interval:
## 0.6837869 1.0000000
## sample estimates:
## p
## 0.875
A few details to unpack:
x argument and the denominator (the sample size for
one-sample data) as the n argument.p
argument.alternative argument is used for a
two-sided test, but we could set it to "less" or
"greater" for one-sided testsprop.test() applies the Yates’ continuity
correction, but we can turn this off using correct = FALSE
correct = FALSE in a real
data analysis, but we are only doing it in this lab to see that the
\(p\)-values from
prop.test() exactly match the ones we calculate
ourselvesQuestion #1: For this question you should use a random sample of \(n=200\) NFL games played between 2018 and 2023.
nfl_sample = read.csv("https://remiller1450.github.io/data/nfl_sample.csv")
game_outcome takes on a value of 1 when the home team won.
Suppose we’d like to test the hypothesis that home teams have an
advantage and are more likely to win than away teams using this sample
of \(n=200\) games. Briefly explain why
this scenario is an example of one-sample categorical data.game_outcome=0.5) are not considered wins.prop.test() to perform the
one-sample \(Z\)-test and confirm that
\(p\)-value you get matches the one you
found using StatKey.\(~\)
The Normal probability model that the one-sample Z-test relies upon is only reasonable when at least 10 observations in the sample belong to each category involved in the proportions. Or, put differently, when \(n\cdot p \geq 10\) and \(n\cdot (1-p) \geq 10\).
The primary issue with the Normal model in these small-sample situations is that the null distribution contains a small number of discrete possibilities that cannot be reliably approximated by a continuous curve. The exact binomial test overcomes this by using the binomial probability distribution to calculate the probability of each discrete outcome present in the null distribution.
The example below uses binom.test() to perform the exact
binomial test on the data from our helper-hinderer example:
binom.test(x = 14, n = 16, p = 0.5, alternative = "greater")
##
## Exact binomial test
##
## data: 14 and 16
## number of successes = 14, number of trials = 16, p-value = 0.00209
## alternative hypothesis: true probability of success is greater than 0.5
## 95 percent confidence interval:
## 0.6561748 1.0000000
## sample estimates:
## probability of success
## 0.875
You should note that this function uses the same arguments/syntax as
prop.test(), but the \(p\)-value we get is slightly different.
You should also notice that we only observed 2 choices of the “hinderer” toy in the study, so the large sample condition of the \(Z\)-test is not met, leading us to prefer the exact binomial test for an analysis of these data.
Question #2:
nfl_sample data and
hypotheses from Question 1, report the \(p\)-value of an exact binomial test.\(~\)
Recall that the one-sample \(T\)-test is performed in almost exactly the same manner as the one-sample \(Z\)-test:
The t.test() function is used to perform the one-sample
\(t\)-test. The R code
below performs this test on a random sample of \(n=200\) patients at the CMU ICU, testing
whether their mean systolic blood pressure differs from 120 mm/Hg.
## Load data
icu_data = read.csv("https://remiller1450.github.io/data/ICUAdmissions.csv")
## One-sample T-test
t.test(x = icu_data$Systolic, mu = 120, alternative = "two.sided")
##
## One Sample t-test
##
## data: icu_data$Systolic
## t = 5.2702, df = 199, p-value = 3.529e-07
## alternative hypothesis: true mean is not equal to 120
## 95 percent confidence interval:
## 127.6852 136.8748
## sample estimates:
## mean of x
## 132.28
Notice that the variable of interest is passed via the x
argument, and null hypothesis is given via the mu
argument.
Question #3: In this question you’ll continue using
the nfl_sample data, this time looking at the
score_diff variable, which records the numerical difference
in points scored, calculated as home team score - away team score.
score_diff variable. State the null hypothesis for this
research question using statistical symbols (ie: \(H_0: \mu = \_\)) and calculate the \(T\)-value corresponding with the observed
sample mean.t.test() to perform the
one-sample \(T\)-test and confirm that
\(p\)-value you get matches the one you
found using StatKey.\(~\)
The Wilcoxon Signed-Rank Test is a non-parametric alternative to the one-sample \(T\)-test for a single quantitative variable. It is often used when the sample size is small and it is unreasonable to assume the data came from a Normally distributed population.
While the \(T\)-test considers whether the average observation is unusually far from the hypothesized value, the Wilcoxon signed-rank test considers whether observations tend to disproportionately fall above or below the hypothesized value using the ranks of the observations. More specifically, the Wilcoxon Signed-Rank Test first ranks each data-point based upon its absolute value, then the data-points are grouped according to their sign (positive or negative). The test then compares the sum of each data-point’s rank multiplied by its sign against a null distribution to produce a \(p\)-value.
We won’t cover the precise details of the test, but its comparison of signed ranks amounts to a test of whether the population median is a certain hypothesized value, or \(H_0: m = \_\).
The code below demonstrates this test for our CMU ICU systolic blood pressure example:
## Signed Rank Test
wilcox.test(x = icu_data$Systolic, mu = 120)
##
## Wilcoxon signed rank test with continuity correction
##
## data: icu_data$Systolic
## V = 13100, p-value = 2.723e-07
## alternative hypothesis: true location is not equal to 120
Question #4: A veterinary anatomist measured the
nerve cell density at two different locations in the intestine, site 1 -
the mid-region of the jejunum and site 2 - the mesenteric region of the
jejunum. The nerve cell density (thousands of cells per mm\(^3\)) was measured in each location, with
the difference recorded as the variable diff, which is the
focus of this analysis.
diff using 15 bins. Based upon
this histogram and the sample size, explain whether you believe the
one-sample \(T\)-test is appropriate
for these data.\(~\)
Question #5: On a previous assignment, you worked with the “TSA claims” data set, a random sample of \(n=5000\) of claims made by travelers against the Transport Security Administration (TSA) between 2003 and 2008, the first five years that the agency existed. In this question you will investigate whether the average paid claim amount is less than $200.
R used to perform the test, provide a one-sentence
conclusion and a one-sentence justification for the choice of test,
being mindful of the assumptions involved.\(~\)
Question #6: On a previous assignment, you worked with “ACS Employment” data set, which is a random sample of 1287 employed individuals collected as part of the American Community Survey (ACS) performed by the US Census Bureau. In this question you will investigate a claim by Forbes that 89% of US adults have health insurance.
R used to perform the test, provide a one-sentence
conclusion and a one-sentence justification for the choice of test,
being mindful of the assumptions involved.