\(~\)
Recall that the \(T\)-distribution was created using the assumption that data are sampled from a Normally distributed population. Thus, for small sample sizes, the \(T\)-test is only appropriate when the sample data show no signs of skewness or extreme outliers.
If the sample data are skewed (or have outliers on one side of the distribution) statisticians may apply a data transformation, with the log-transformation being the most common choice (for right-skewed data):
| Original Value | Log2(Original Value) |
|---|---|
| 1 | 0.000000 |
| 2 | 1.000000 |
| 4 | 2.000000 |
| 10 | 3.321928 |
| 100 | 6.643856 |
Log-transformations change the scale on which the data are measured, such that the base of the logarithm determines the multiplicative change corresponding to a one-unit increase on the transformed scale:
Note that:
The example below uses the TSA claims data to illustrate the impact of a log10 transformation on a highly skewed distribution:
## Let's load some right-skewed data
library(dplyr)
tsa_eyeglasses <- read.csv("https://remiller1450.github.io/data/tsa.csv") %>%
filter(Item == "Eyeglasses - (including contact lenses)")
## We use the log() function to create a new log-transformed variable
tsa_eyeglasses$log10_claim = log(tsa_eyeglasses$Claim_Amount, base = 10)
## Graph the distributions of claim amount and log10(claim amount)
ggplot(data = tsa_eyeglasses, aes(x = Claim_Amount)) + geom_histogram(bins = 30) + theme_light()
ggplot(data = tsa_eyeglasses, aes(x = log10_claim)) + geom_histogram(bins = 30) + theme_light()
\(~\)
A key property of logarithms is that differences on the log-scale are ratios (relative changes) on the original scale when the transformation is undone: \[log_{10}(X) - log_{10}(Y) = log_{10}(X/Y) \\ 10^{log_{10}(X/Y)} = X/Y\]
So, a two-sample \(t\)-test using log-transformed data is actually testing a ratio of means, which is often more easily interpretable than the absolute difference in means for many applications.
For example, if the estimated difference in means on the log10-scale is \(0.20\), using exponentiation to undo the transformation gives \(10^{0.20}=1.585\), indicating the mean of the first group is estimated to be approximately 58.5% larger than the mean of the other group.
Few final detail to note:
\(~\)
A 2010 study aimed to link recreational drug use with risky driving
behavior, such as following too closely, using data collected in an
advanced driving simulator. Study participants had to follow a lead
vehicle that was programmed to erratically vary its speed. The
researchers recorded each participant’s average following distance
behind this vehicle as the study outcome, which is recorded as the
variable D in the data set.
tailgating = read.csv("https://remiller1450.github.io/data/Tailgating.csv")
The “tailgating” data set described above contains several large outliers that are easily identified when graphing the distribution following distances in each group:
library(ggplot2)
ggplot(tailgating, aes(x = D, y = Drug)) + geom_boxplot()
Our analysis will compare participants in the hardest drug group
(MDMA, more commonly known as ecstasy or Molly) and second
hardest drug group (THC, the primary psychoactive component
of cannabis).
Question #1: For this question, use the subset
created below that uses the filter() function (introduced
in Lab 4) to keep only cases from the MDMA and THC groups:
tailgating_subset = filter(tailgating, Drug %in% c("MDMA", "THC"))
table() function to
verify the sample sizes of the MDMA and THC groups as \(n_1=16\) and \(n_2=39\), respectively.R to perform a two-sample \(T\)-test comparing the mean following
distances of drivers who recreationally use MDMA or THC. Report a
1-sentence conclusion that includes context, the strength of evidence,
and the observed sample means (recognizing that you can get them from
the t.test() output).\(~\)
Recall that the THC group includes an extreme outlier whose mean following distance is over 300 feet. You can use the code below to create a subset that removes this individual from the dataset:
tailgating_no_outliers = filter(tailgating_subset, D < 100)
Before proceeding to analyze the data with the outlier, you should carefully consider whether it is justifiable to remove these individuals:
Question #2: For this question you should use the
tailgating_no_outliers subset created above.
\(~\)
The variable LD in the tailgating data is the natural
logarithm of D, the participant’s average following
distance. While this variable was already provided to you in the data,
you should recognize that you could create a log-transformed version of
a variable yourself using the log() function:
tailgating_subset$new_log_distance = log(tailgating_subset$D)
Notice how this transformation makes the outliers and rightward skew present in the non-transformed data noticeably less prominent:
ggplot(tailgating_subset, aes(x = LD, y = Drug)) + geom_boxplot()
Question #3: For this question you should use the
tailgating_subset data that contains outliers.
LD. Be careful to use the
entire data set, not the version you used in Question 2 that
had the outlier removed. Briefly explain how the results of this test
compare to those from Question 1 and Question 2.LD in the
THC group was 3.55, while the mean in the MDMA group was 3.28, a
difference of 0.27 (on the log-scale). Use this information to estimate
the relative difference in following distances by undoing the
log-transformation. Write a 1-sentence conclusion to your hypothesis
test from Part A that incorporates this relative difference.