\(~\)
Directions:
- Homework must be completed individually. Any guidance or help
received from mentors, classmates, online resources other than course
materials (including AI/LLMs) must be acknowledged.
- Clearly organize your responses so that each question (1, 2, etc.)
and sub-question (Part A, B, C, etc.) are clearly identifiable.
- Please submit a single Jupyter notebook (.ipynb file) displaying all
output and recording textual answers using neatly formatted markdown
chunks.
- Your submission should be made via Canvas no later than 11:59pm on
the assigned due-date.
Question #1 (PCA and Data handling)
In several of our labs you’ve worked with the MNIST dataset, a
collection of \(n=6000\) handwritten
digits randomly sampled from the MNIST database.
Recall that the data you’ve been given was flattened, meaning each row
represents a single image with 784 columns used to represent grayscale
intensities of the 28 by 28 pixels in the image.
https://remiller1450.github.io/data/mnist_small.csv
- Part A: Read the data into Python then separate the
pixel intensities from the label column, thereby creating the objects
pixels and labels. Then separate the data into training and testing sets
in an 80-20 ratio using
random_state=7.
- Part B: Create a version of the training data that
is a 3-dimensional
numpy array where the first axis (axis
0) is the number of samples. Then, use io.imshow() with the
argument cmap = plt.cm.Greys (assuming you loaded the
matplotlib library under the commonly used alias “plt”) to display the
6th sample in the training set.
- Part C: Fit a PCA model using
n_components = 10 using the flattened training pixels.
Next, use the inverse_transform() method and reshape the
result to match the dimensions of the 3-dimensional numpy array you
created in Part B.
- Part D: Replace any negative values in the results
from Part C, then repeat the steps of Part B to visualize the 6th sample
after dimension reduction via PCA. Briefly describe how these images
visually compare, and how they are related.
- Part E: The original data set used 4,704,000
numeric values to store the pixel intensities of the entire set of
images. If PCA with
n_components = 10 is used to “compress”
these data, how many numerical values need to be stored? Hint:
Think about both the dimensions of the data after using PCA for
dimension reduction as well as the number of values that are needed to
reconstruct the original shape.
\(~\)
Question #2 (Cross-validation)
Consider a toy data set of \(n=200\)
observations generated via \(Y = f(X) +
\epsilon\) where \(f(X) = 2x_1 +
10\) and \(\epsilon \sim N(0,\sigma
=5)\). Notice that the true relationship between \(x_1\) and \(Y\) is linear, with 5 units of irreducible
error. These data can be found at the URL given below:
https://remiller1450.github.io/data/toy_linear_data.csv
For the questions that follow you should use all \(n=200\) observations and do not need to
perform a train-test split.
- Part A: In your own words, explain whether a KNN
regressor or decision tree model is better suited to estimating the true
\(f()\). You should rely upon
conceptual arguments in favor of your chosen method, not empirical
investigations using the provided data.
- Part B: Use a for loop to create your own
implementation of 4-fold cross-validation for a decision tree model with
a maximum depth of 4. Randomly assign each observation to one of four
folds using
random_state=7. You are encouraged to consult
the pseudocode from our lecture slides and the
np.random.choice() function. Use your implementation to
report the cross-validated RMSE.
- Part C: Replace the decision tree model in your
cross validation loop with a
LinearRegression model from
the linear_model library of sklearn. How does
the change in model impact the cross-validated RMSE?
- Part D: The decision tree model in Part B used a
maximum depth of 4. If you performed a grid search allowing this
hyperparameter to be any positive integer, would you expect to find one
that produces a cross-validated RMSE as low as the linear regression
modeling approach in Part C? Briefly explain. Note: you should
make a conceptual argument and should not actually perform this
grid search.
- Part E: In your own words, briefly explain why
cross-validation is preferable to using the RMSE on the training data to
choose between the two competing models described in Parts B and C.
\(~\)
Question #3 (Application)
For this question you should use the dataset available here:
https://remiller1450.github.io/data/beans.csv
This data set was constructed using a computer vision system that
segmented 13,611 images of 7 types of dry beans, extracting 16 features
(12 dimensional features, 4 shape-form features). Additional details are
contained in
this paper
The following questions should be answered using Python code,
including functions in the sklearn library. Unless
otherwise indicated, use the 'accuracy' scoring criteria
during model tuning and evaluation.
- Part A: Read these data into Python and perform a
90-10 training-testing split using
random_state=1.
- Part B: Separate the outcome, “Class”, from the
predictors and graph a histogram of every predictor. Based on both the
predictor distributions and the models under consideration, do you think
that re-scaling and/or transformation should be part of a data
preparation pipeline? You should note that \(k\)-nearest neighbors and decision trees
will both be considered.
- Part C: Use the
corrcoef() function in
numpy to explore the pairwise correlations between
predictors. Based upon these correlations, do you think dimension
reduction via principal component analysis should be considered as part
of your data preparation pipeline?
- Part D: Create a machine learning pipeline that
includes the data preparation steps you deemed important in Parts B and
C. Then, perform a grid search using 5-fold cross-validation to find a
well-fitting \(k\)-nearest neighbors
model. Your search should explore at least two variations of
your data preparation steps, at least three values of \(k\), both euclidean or Manhattan
distance, and both uniform or distance weighting. Your data
preparation variations must include either different scaling methods,
different PCA configurations, or both. Report both the best
hyperparameter configuration and its mean cross-validated accuracy for
the best performing approach.
- Part E: Repeat the same basic steps of Part D to
find a well-fitting decision tree model. Your search should
explore at least two variations of your data preparation steps
(ie: two different scalers, or two different numbers of retained
principal components), at least three different maximum depths,
and at least two values of minimum samples required to split a
node. Your data preparation variations may include the possibility that
a preprocessing step is skipped using
'passthrough'. Report
both the best hyperparameter configuration and its mean cross-validated
accuracy for the best performing approach.
- Part F: Construct a final search using
GridSearchCV() that compares the best KNN and best decision
tree approaches identified in Parts D and E using a common set of 5-fold
cross-validation splits. Print the overall best estimator.
- Part G: Create a visualization of the confusion
matrix for your best estimator from Part F that displays
classification results for the test data. Report the most common
type of misclassification made by the model.
- Part H: Report both the macro-averaged and
micro-averaged F1-scores of the best classifier from Part F on the
test data. Which of these approaches (macro or micro averaging) do
you believe is more appropriate for this application? Or are both
approaches reasonable? Briefly explain.
\(~\)