\(~\)

Onboarding

The main theme of this lab is answering the question “which one-sample test should I use?”

As a general decision framework, before running any statistical test we should consider the following:

  1. The type of data (categorical or quantitative)
  2. The sample size
  3. Whether the assumptions of the test appear reasonable
Outcome Type Sample Characteristics Recommended Test
Categorical Large sample (at least 10 observations in each category) One-sample Z-test - prop.test()
Categorical Small sample (fewer than 10 observations in one category) Exact Binomial Test - binom.test()
Quantitative Population approximately Normal OR \(n \geq30\) One-sample T-test - t.test()
Quantitative Small sample (\(n<30\)) and strong evidence of non-Normality Wilcoxon Signed-Rank Test - wilcox.test()

We have covered the one-sample Z-test and T-test in our previous two lectures, and this lab will briefly cover the exact binomial and Wilcoxon signed-rank tests.

After introducing the R implementation of each of these tests using examples and practice questions, the lab features two questions to practice applying a consistent statistical workflow:

  1. Identifying the variable involved in the research question and generating appropriate data visualizations and summary statistics.
  2. Choosing an appropriate hypothesis test and providing a justification
  3. Writing a proper one-sentence conclusion based upon the results of the test.

\(~\)

Lab

One-sample \(Z\)-Test

Recall the one-sample \(Z\)-test involves two primary steps:

  1. Calculate \(Z=\frac{\hat{p} - p}{\sqrt{p(1-p)/n}}\) using \(p\) from the null hypothesis and \(\hat{p}\) from the sample
  2. Compare \(Z\) to the N(0,1) distribution to calculate the \(p\)-value

The prop.test() function in R can be used to find the \(p\)-value produced by the one-sample Z-test. The code below uses prop.test() to replicate the one-sample \(Z\)-test on the infant toy choice data from our previous lecture:

prop.test(x = 14, n = 16, p = 0.5, alternative = "greater", correct = FALSE)
## 
##  1-sample proportions test without continuity correction
## 
## data:  14 out of 16, null probability 0.5
## X-squared = 9, df = 1, p-value = 0.00135
## alternative hypothesis: true p is greater than 0.5
## 95 percent confidence interval:
##  0.6837869 1.0000000
## sample estimates:
##     p 
## 0.875

A few details to unpack:

  • We provide the numerator of the sample proportion as the x argument and the denominator (the sample size for one-sample data) as the n argument.
  • The null hypothesis is specified by the p argument.
  • By default the alternative argument is used for a two-sided test, but we could set it to "less" or "greater" for one-sided tests
  • By default prop.test() applies the Yates’ continuity correction, but we can turn this off using correct = FALSE
    • Note you typically won’t set correct = FALSE in a real data analysis, but we are only doing it in this lab to see that the \(p\)-values from prop.test() exactly match the ones we calculate ourselves

Question #1: For this question you should use a random sample of \(n=200\) NFL games played between 2018 and 2023.

nfl_sample = read.csv("https://remiller1450.github.io/data/nfl_sample.csv")
  • Part A: In these data, the variable game_outcome takes on a value of 1 when the home team won. Suppose we’d like to test the hypothesis that home teams have an advantage and are more likely to win than away teams using this sample of \(n=200\) games. Briefly explain why this scenario is an example of one-sample categorical data.
  • Part B: State the null hypothesis for the research question Part A using statistical symbols (ie: \(H_0: p = \_\)) and calculate the \(Z\)-value corresponding with the observed sample proportion. For the purposes of this question you can assume ties (game_outcome=0.5) are not considered wins.
  • Part C: Use the Normal Probability curve menu of StatKey to find a one-sided \(p\)-value for the one-sample Z-test described in Parts A and B. Report this \(p\)-value as part of a proper one-sentence conclusion. Remember that you should include context, strength of evidence, and direction of the effect (if one was found).
  • Part D: Use prop.test() to perform the one-sample \(Z\)-test and confirm that \(p\)-value you get matches the one you found using StatKey.

\(~\)

Exact Binomial Test

The Normal probability model that the one-sample Z-test relies upon is only reasonable when at least 10 observations in the sample belong to each category involved in the proportions. Or, put differently, when \(n\cdot p \geq 10\) and \(n\cdot (1-p) \geq 10\).

The primary issue with the Normal model in these small-sample situations is that the null distribution contains a small number of discrete possibilities that cannot be reliably approximated by a continuous curve. The exact binomial test overcomes this by using the binomial probability distribution to calculate the probability of each discrete outcome present in the null distribution.

The example below uses binom.test() to perform the exact binomial test on the data from our helper-hinderer example:

binom.test(x = 14, n = 16, p = 0.5, alternative = "greater")
## 
##  Exact binomial test
## 
## data:  14 and 16
## number of successes = 14, number of trials = 16, p-value = 0.00209
## alternative hypothesis: true probability of success is greater than 0.5
## 95 percent confidence interval:
##  0.6561748 1.0000000
## sample estimates:
## probability of success 
##                  0.875

You should note that this function uses the same arguments/syntax as prop.test(), but the \(p\)-value we get is slightly different.

You should also notice that we only observed 2 choices of the “hinderer” toy in the study, so the large sample condition of the \(Z\)-test is not met, leading us to prefer the exact binomial test for an analysis of these data.

Question #2:

  • Part A: Using the nfl_sample data and hypotheses from Question 1, report the \(p\)-value of an exact binomial test.
  • Part B: Considering the assumptions of the one-sample \(Z\)-test, is it necessary to use an exact binomial test for these data in order to obtain a reliable estimate of the \(p\)-value? Briefly explain.

\(~\)

One-sample \(T\)-test

Recall that the one-sample \(T\)-test is performed in almost exactly the same manner as the one-sample \(Z\)-test:

  1. Calculate \(T=\frac{\overline{x} - \mu}{s/\sqrt{n}}\) using \(\mu\) from the null hypothesis and \(\overline{x}\) and \(s\) from the sample
  2. Compare \(T\) to the \(t\)-distribution with \(n-1\) degrees of freedom to calculate the \(p\)-value

The t.test() function is used to perform the one-sample \(t\)-test. The R code below performs this test on a random sample of \(n=200\) patients at the CMU ICU, testing whether their mean systolic blood pressure differs from 120 mm/Hg.

## Load data
icu_data = read.csv("https://remiller1450.github.io/data/ICUAdmissions.csv")

## One-sample T-test
t.test(x = icu_data$Systolic, mu = 120, alternative = "two.sided")
## 
##  One Sample t-test
## 
## data:  icu_data$Systolic
## t = 5.2702, df = 199, p-value = 3.529e-07
## alternative hypothesis: true mean is not equal to 120
## 95 percent confidence interval:
##  127.6852 136.8748
## sample estimates:
## mean of x 
##    132.28

Notice that the variable of interest is passed via the x argument, and null hypothesis is given via the mu argument.

Question #3: In this question you’ll continue using the nfl_sample data, this time looking at the score_diff variable, which records the numerical difference in points scored, calculated as home team score - away team score.

  • Part A: Suppose we’d like to test the hypothesis that home teams are advantaged using these data and the score_diff variable. State the null hypothesis for this research question using statistical symbols (ie: \(H_0: \mu = \_\)) and calculate the \(T\)-value corresponding with the observed sample mean.
  • Part B: Use the t-distribution menu of StatKey to calculate a one-sided \(p\)-value for the hypothesis described in Part A. Report this \(p\)-value as part of a proper one-sentence conclusion. Remember that you should include context, strength of evidence, and direction of the effect (if one was found).
  • Part C: Use t.test() to perform the one-sample \(T\)-test and confirm that \(p\)-value you get matches the one you found using StatKey.

\(~\)

Wilcoxon Signed-Rank Test

The Wilcoxon Signed-Rank Test is a non-parametric alternative to the one-sample \(T\)-test for a single quantitative variable. It is often used when the sample size is small and it is unreasonable to assume the data came from a Normally distributed population.

While the \(T\)-test considers whether the average observation is unusually far from the hypothesized value, the Wilcoxon signed-rank test considers whether observations tend to disproportionately fall above or below the hypothesized value using the ranks of the observations. More specifically, the Wilcoxon Signed-Rank Test first ranks each data-point based upon its absolute value, then the data-points are grouped according to their sign (positive or negative). The test then compares the sum of each data-point’s rank multiplied by its sign against a null distribution to produce a \(p\)-value.

We won’t cover the precise details of the test, but its comparison of signed ranks amounts to a test of whether the population median is a certain hypothesized value, or \(H_0: m = \_\).

The code below demonstrates this test for our CMU ICU systolic blood pressure example:

## Signed Rank Test
wilcox.test(x = icu_data$Systolic, mu = 120)
## 
##  Wilcoxon signed rank test with continuity correction
## 
## data:  icu_data$Systolic
## V = 13100, p-value = 2.723e-07
## alternative hypothesis: true location is not equal to 120

Question #4: A veterinary anatomist measured the nerve cell density at two different locations in the intestine, site 1 - the mid-region of the jejunum and site 2 - the mesenteric region of the jejunum. The nerve cell density (thousands of cells per mm\(^3\)) was measured in each location, with the difference recorded as the variable diff, which is the focus of this analysis.

https://remiller1450.github.io/data/horse_nerves.csv

  • Part A: Create a histogram displaying the distribution of the variable diff using 15 bins. Based upon this histogram and the sample size, explain whether you believe the one-sample \(T\)-test is appropriate for these data.
  • Part B: Perform a Wilcoxon Signed-Rank Test to assess whether the population-level difference in nerve cell density across these two locations is zero. Report the two-sided \(p\)-value and a one-sentence conclusion.
  • Part C: Now perform a one-sample \(T\)-test to evaluate the same hypotheses used in Part B. Considering the \(p\)-value of this test, why was the assumption check you performed in Part A important? Briefly explain.

\(~\)

Practice (required)

Question #5: On a previous assignment, you worked with the “TSA claims” data set, a random sample of \(n=5000\) of claims made by travelers against the Transport Security Administration (TSA) between 2003 and 2008, the first five years that the agency existed. In this question you will investigate whether the average paid claim amount is less than $200.

https://remiller1450.github.io/data/tsa_small.csv

  • Part A: Create an appropriate data visualization for the data involved in the research question described above.
  • Part B: Provide appropriate descriptive statistics that summarize the key aspects of the data visualization you created in Part A.
  • Part C: Perform an appropriate hypothesis test addressing the research question described above. In addition to the R used to perform the test, provide a one-sentence conclusion and a one-sentence justification for the choice of test, being mindful of the assumptions involved.

\(~\)

Question #6: On a previous assignment, you worked with “ACS Employment” data set, which is a random sample of 1287 employed individuals collected as part of the American Community Survey (ACS) performed by the US Census Bureau. In this question you will investigate a claim by Forbes that 89% of US adults have health insurance.

https://remiller1450.github.io/data/EmployedACS.csv

  • Part A: Create an appropriate data visualization for the data involved in the research question described above.
  • Part B: Provide appropriate descriptive statistics that summarize the key aspects of the data visualization you created in Part A.
  • Part C: Perform an appropriate hypothesis test addressing the research question described above. In addition to the R used to perform the test, provide a one-sentence conclusion and a one-sentence justification for the choice of test, being mindful of the assumptions involved.