Directions:
- Submit your final answers and all supporting work on Canvas (under
the “Assignments” tab)
- You may submit any file format for this assignment
- Homework is intended to be individual work. While you may
discuss the assignment with your peers, you should submit answers that
are uniquely your own.
- Any assistance you receive from resources/materials not on our
course website (such as other websites, course mentors, peer students,
AI, etc.) should be clearly acknowledged
Question #1 (Sampling)
The following are conceptual questions related to sampling. Please
answer each in 1-2 sentences.
- Part A: A marine biologist would like to estimate
the average size of a certain species of fish in a lake. They randomly
place nets in ten different locations and use the fish collected by
these nets as their sample. Will this sample be representative of all
fish of this species in the lake, or will it be biased? If it is biased,
will it overestimate or underestimate the true value?
- Part B: Suppose we want to estimate household size,
where a “household” is defined as people living together in the same
dwelling, and sharing living accommodations. If we select students at
random at an elementary school and ask them what their family size is,
will this be a good measure of household size? Or will our average be
biased? If it is biased, will it overestimate or underestimate the true
value?
- Part C: A statistics student plans on estimating
the average word length in speeches written by Abraham Lincoln using a
sample of words from the Gettysburg Address. They print off a copy of
the speech, hang it somewhere in Noyce, then throw ten darts at the
speech and record the ten words closest to where the darts land. Do you
believe the lengths of these words will be a good measure of the actual
average word length in Lincoln’s speeches? If not, will they
overestimate or underestimate the true value?
- Part D: A company executive would like to gauge the
satisfaction of employees. They randomly select 30 workers from a
comprehensive company database and send them an anonymous questionnaire.
If all of the selected workers return the questionnaire, will the
responses be representative of all employees at the company? If not,
will they overestimate or underestimate the true levels of employee
satisfaction?
\(~\)
Question #2 (One-Sample Hypothesis Testing)
For this question you should use the data found at the link below.
These data are a random sample of 1287 employed individuals collected as
part of the American Community Survey (ACS) performed by the US Census
Bureau.
https://remiller1450.github.io/data/EmployedACS.csv
- Part A Consider an analysis of the variable
USCitizen, which records whether the respondent holds US
citizenship (a value of 1) or is not a US citizen (a value of 0). Is
this an example of one-sample categorical or one-sample quantitative
data? You do not need to explain your answer.
- Part B: Create an appropriate data visualization
showing the distribution of the variable
USCitizen in the
sample.
- Part C: According
to Ballotpedia 7.1% of the US population were non-citizens as of
2014. Suppose you hypothesize that the true percentage of non-citizens
differs from the percentage reported by Ballotpedia and would like to
use these data as evidence. Using appropriate statistical notation,
write the null and alternative hypotheses for how you could use these
data to perform a statistical test evaluating your hypothesis.
- Part D: To evaluate the hypotheses in Part C,
should you use a Z test or a T test? Briefly explain your answer.
- Part E: Are the sample size conditions met to use
the test you identified in Part D? Briefly address these
conditions.
- Part F: Calculate the test statistic (Z or T score)
for the hypothesis test you identified in Parts C-D. Show your
work.
- Part G: Use either
prop.test() or
t.test() (whichever is appropriate for the test you
identified) to find a two-sided \(p\)-value. Report your \(p\)-value as part of a one-sentence
conclusion. Make sure you include all of the components of a proper
conclusion (context, strength of evidence, and type of
relationship).
- Part H: Use probability notation (ie: \(Pr(\ldots)\)) and statistical symbols (ie:
\(p\), \(\mu\), etc.) to define the \(p\)-value you calculated in Part G. It is
okay to provide the probability statement for the one-sided \(p\)-value multiplied by two (ie: \(2\cdot Pr(\ldots)\))
\(~\)
Question #3 (Descriptive Statistics and Data Visualizations)
For this question you should continue using the ACS employed
individuals survey data provided as part of Question #2.
- Part A: Using only the contents of Labs 2 and 3,
construct an appropriate data visualization for each of the following
variables:
- Part B: Using your visualization from Part A, as
well as additional descriptive statistics, describe the shape,
center, and spread of the distribution of the
average weekly incomes reported by individuals in the sample.
If the distribution appears highly skewed, you should report measures of
center and spread that are robust to the influence of outliers.
- Part C: Using your visualization from Part A, as
well as additional descriptive statistics, describe the shape,
center, and spread of the distribution of the
average weekly hours worked reported by individuals in the
sample. If the distribution appears highly skewed, you should report
measures of center and spread that are robust to the influence of
outliers.