\(~\)

Onboarding

Recall that the \(T\)-distribution was created using the assumption that data are sampled from a Normally distributed population. Thus, for small sample sizes, the \(T\)-test is only appropriate when the sample data show no signs of skewness or extreme outliers.

If the sample data are skewed (or have outliers on one side of the distribution) statisticians may apply a data transformation, with the log-transformation being the most common choice (for right-skewed data):

Original Value Log2(Original Value)
1 0.000000
2 1.000000
4 2.000000
10 3.321928
100 6.643856

Log-transformations change the scale on which the data are measured, such that the base of the logarithm determines the multiplicative change corresponding to a one-unit increase on the transformed scale:

Note that:

  1. On the original scale (non-transformed data) the values 1, 2, 4, and 10 are closely bunched together, and 100 is a clear outlier. But after log-transformation, these closely bunched values become spread apart, and 100 isn’t as much of an outlier.
  2. One the log-2 scale, the distance between 1 and 2 is the same as the distance between 2 and 4, as both of these are a 2-fold increase (doubling).
  3. On the log-10-scale, the distance between 1 and 10 is the same as the distance between 10 and 100, as both of these are a 10-fold increase.

The example below uses the TSA claims data to illustrate the impact of a log10 transformation on a highly skewed distribution:

## Let's load some right-skewed data
library(dplyr)
tsa_eyeglasses <- read.csv("https://remiller1450.github.io/data/tsa.csv") %>% 
  filter(Item == "Eyeglasses - (including contact lenses)")

## We use the log() function to create a new log-transformed variable
tsa_eyeglasses$log10_claim = log(tsa_eyeglasses$Claim_Amount, base = 10)

## Graph the distributions of claim amount and log10(claim amount)
ggplot(data = tsa_eyeglasses, aes(x = Claim_Amount)) + geom_histogram(bins = 30) + theme_light()
ggplot(data = tsa_eyeglasses, aes(x = log10_claim)) + geom_histogram(bins = 30) + theme_light()

\(~\)

Interpreting Log-Transformed Data

A key property of logarithms is that differences on the log-scale are ratios (relative changes) on the original scale when the transformation is undone: \[log_{10}(X) - log_{10}(Y) = log_{10}(X/Y) \\ 10^{log_{10}(X/Y)} = X/Y\]

So, a two-sample \(t\)-test using log-transformed data is actually testing a ratio of means, which is often more easily interpretable than the absolute difference in means for many applications.

For example, if the estimated difference in means on the log10-scale is \(0.20\), using exponentiation to undo the transformation gives \(10^{0.20}=1.585\), indicating the mean of the first group is estimated to be approximately 58.5% larger than the mean of the other group.

Few final detail to note:

  1. Statisticians use \(log()\) without a subscript specifying the base to mean the natural logarithm (a base of \(e\)).
  2. Because \(\sum_{i=1}^{n}log(x_i)/n \ne log(\sum_{i=1}^{n}x_i/n)\) we technically get a ratio of geometric means (rather than arithmetic means) after undoing a log-transformation. For our purposes, this distinction is not particularly important. The main idea is that by analyzing log-transformed data we can assess relative changes between groups.

\(~\)

Lab

Application - Drug Use and Tailgating

A 2010 study aimed to link recreational drug use with risky driving behavior, such as following too closely, using data collected in an advanced driving simulator. Study participants had to follow a lead vehicle that was programmed to erratically vary its speed. The researchers recorded each participant’s average following distance behind this vehicle as the study outcome, which is recorded as the variable D in the data set.

tailgating = read.csv("https://remiller1450.github.io/data/Tailgating.csv")

The “tailgating” data set described above contains several large outliers that are easily identified when graphing the distribution following distances in each group:

library(ggplot2)
ggplot(tailgating, aes(x = D, y = Drug)) + geom_boxplot()

Our analysis will compare participants in the hardest drug group (MDMA, more commonly known as ecstasy or Molly) and second hardest drug group (THC, the primary psychoactive component of cannabis).

Question #1: For this question, use the subset created below that uses the filter() function (introduced in Lab 4) to keep only cases from the MDMA and THC groups:

tailgating_subset = filter(tailgating, Drug %in% c("MDMA", "THC"))
  • Part A: Use the table() function to verify the sample sizes of the MDMA and THC groups as \(n_1=16\) and \(n_2=39\), respectively.
  • Part B: Considering the group sizes and the distribution of the outcome variable, would it be appropriate to analyze these data using a two-sample \(T\)-test? Briefly explain why this test is or isn’t appropriate for these data.
  • Part C: Temporarily ignoring your response to Part B, use R to perform a two-sample \(T\)-test comparing the mean following distances of drivers who recreationally use MDMA or THC. Report a 1-sentence conclusion that includes context, the strength of evidence, and the observed sample means (recognizing that you can get them from the t.test() output).

\(~\)

Outliers

Recall that the THC group includes an extreme outlier whose mean following distance is over 300 feet. You can use the code below to create a subset that removes this individual from the dataset:

tailgating_no_outliers = filter(tailgating_subset, D < 100)

Before proceeding to analyze the data with the outlier, you should carefully consider whether it is justifiable to remove these individuals:

  • On one hand, selectively disregarding data will introduce bias into the analysis.
  • On the other hand, they might not have understood the study instructions or taken the driving simulation seriously, so including them also might unfairly influence the analysis.

Question #2: For this question you should use the tailgating_no_outliers subset created above.

  • Part A: Consider the two arguments presented immediately prior to this question. Which do you find more compelling? Briefly explain. Note: There is not necessarily a right or wrong answer here.
  • Part B: Repeat the two-sample \(T\)-test you performed in Question 1 after outliers have been removed. What effect did the outlier have on the \(p\)-value of the test? Does removing the outlier change the decision at the \(\alpha = 0.05\) siginficance threshold?
  • Part C: Notice the sample means of each group are actually closer together after the extreme outlier was removed from the THC group. Considering the formula of the test statistic used in the \(T\)-test, explain how the change in \(p\)-value you observed in Part B is possible given the sample means are now closer to each other.

\(~\)

Log-transformation

The variable LD in the tailgating data is the natural logarithm of D, the participant’s average following distance. While this variable was already provided to you in the data, you should recognize that you could create a log-transformed version of a variable yourself using the log() function:

tailgating_subset$new_log_distance = log(tailgating_subset$D) 

Notice how this transformation makes the outliers and rightward skew present in the non-transformed data noticeably less prominent:

ggplot(tailgating_subset, aes(x = LD, y = Drug)) + geom_boxplot()

Question #3: For this question you should use the tailgating_subset data that contains outliers.

  • Part A: Repeat the two-sample \(T\)-test from Questions 1 a final time using the log-transformed outcome LD. Be careful to use the entire data set, not the version you used in Question 2 that had the outlier removed. Briefly explain how the results of this test compare to those from Question 1 and Question 2.
  • Part B: The mean value of LD in the THC group was 3.55, while the mean in the MDMA group was 3.28, a difference of 0.27 (on the log-scale). Use this information to estimate the relative difference in following distances by undoing the log-transformation. Write a 1-sentence conclusion to your hypothesis test from Part A that incorporates this relative difference.