Directions:
- Submit your final answers and all supporting work on Canvas (under
the “Assignments” tab)
- You must submit HTML output generated using R Markdown for this
assignment.
- Homework is intended to be individual work. While you may
discuss the assignment with your peers, you should submit answers that
are uniquely your own.
- Any assistance you receive from resources/materials not on our
course website (such as other websites, course mentors, peer students,
AI, etc.) should be clearly acknowledged
Question #1
In Lab 4 you worked with the nfl_sample dataset, which
contained the outcomes of \(n=200\)
randomly chosen NFL games played between 2018 and 2023. The
R below loads these data and creates a new variable that
records the combined score (home team score + away team score) for each
game.
## Load data and create variable of interest
nfl_sample = read.csv("https://remiller1450.github.io/data/nfl_sample.csv")
nfl_sample$combined_score = nfl_sample$home_score + nfl_sample$away_score
- Part A: Across the entire decade of the 1990s, the
average combined score of NFL games was 40.4 points, which can be
considered a known benchmark. Consider the claim that NFL games played
between 2018 and 2023 have higher combined scores than games played in
the 1990s. Using statistical symbols, state an appropriate null and
alternative hypothesis to evaluate this claim. Additionally, report the
relevant sample statistic that will serve as evidence against the null
hypothesis (ie: sample mean, sample proportion, etc.)
- Part B: Consider the type of data (categorical
vs. quantitative) and the characteristics of the sample (sample size,
distributional skewness, etc.) and decide upon a statistical testing
procedure (ie: one-sample Z-test, exact binomial test, etc.) Provide a
brief justification for your testing procedure. Be specific about why
the assumptions of your chosen procedure are or are not reasonable.
- Part C: Use
R to perform the testing
procedure you identified in Part B. Provide a 1-sentence conclusion that
includes context, strength of evidence, and direction of the effect (if
one was found).
\(~\)
Question #2
A regional US bank hypothesized that including photographs of people
on its webpages would increase the rate at which prospective customers
would complete a contact form to start the process of creating an
account or requesting a loan with the bank, an outcome known as
conversion. The bank ran a two-week test where webpage visitors
were randomly routed to one of two landing pages, one featuring a
smiling man or one showing a travel bag on a beach. Of the 3300 people
routed to the page with the smiling man, 55 converted, while of the 3280
people routed to the page with the travel bag, only 44 converted.
- Part A: In your own words, explain whether this is
one-sample or two-sample data, and whether the outcome is categorical or
quantitative.
- Part B: Using proper statistical notation, state
appropriate null and alternative hypotheses to evaluate whether the
bank’s use of a person on its webpage influences conversion.
Additionally, report the observed sample statistic(s) that will serve as
evidence against this hypothesis.
- Part C: Calculate an appropriate test statistic (Z
or T) for the hypothesis test identified in Part B. Show your work,
including each component of the test statistic (observed value, standard
error, etc.) as well as the final value of your test statistic.
- Part D: Use StatKey to compare your test statistic
to the appropriate probability model to obtain a two-sided \(p\)-value. Using this \(p\)-value, provide a 1-sentence conclusion
that includes context, strength of evidence, and direction of the effect
(if one was found).
- Part E: Suppose the bank had run the test for
another week, which would increase the sample size in each group by
roughly 1000. Would you expect this to increase,
decrease, or have no effect on the \(p\)-value? Briefly explain your
reasoning.
\(~\)
Question #3
Researchers at Harvard Medical School conducted an experiment where
infants born with congenital heart defects were randomly assigned to
receive one of two different open-heart surgeries, low-flow bypass or
circulatory arrest. Two years after the surgery, researchers followed up
on each infant and assessed their development, measured by their MDI
(mental development index) and PDI (psychomotor development index)
scores. The data from this experiment are loaded into R by
the code provided below:
ih_data = read.csv("https://remiller1450.github.io/data/InfantHeart.csv")
- Part A: Consider the MDI outcome. In your own
words, explain whether this is one-sample or two-sample data, and
whether the outcome is categorical or quantitative.
- Part B: Create an appropriate data visualization
showing the relationship between MDI and the type of surgery assigned.
Briefly describe what you see in this visualization. Be sure that your
visualization clearly allows comparison of the MDI distributions between
the two surgery groups.
- Part C: Consider the type of data (categorical
vs. quantitative) and the characteristics of the sample (sample size,
distributional skewness, etc.) and decide upon a statistical testing
procedure (ie: Z-test, exact binomial test, etc.) Provide a brief
justification for your testing procedure.
- Part D: Use
R to perform the testing
procedure you identified in Part C. Provide a 1-sentence conclusion that
includes context, strength of evidence, and direction of the effect (if
one was found).
- Part E: Suppose there is measurement error present
in the MDI scores of all of the infants who participated in the study,
but this measurement error can be reduced by having the infants repeat
the assessment on different days and averaging their scores. If the
study protocol implements this change and as a result the sample
standard deviation of MDI scores decreases, would you expect the \(p\)-value to increase,
decrease, or remain unchanged if you re-ran the test
you performed in Part D? Briefly explain your reasoning.