\(~\)

Directions:

Question #1 (PCA and Data handling)

In several of our labs you’ve worked with the MNIST dataset, a collection of \(n=6000\) handwritten digits randomly sampled from the MNIST database. Recall that the data you’ve been given was flattened, meaning each row represents a single image with 784 columns used to represent grayscale intensities of the 28 by 28 pixels in the image.

https://remiller1450.github.io/data/mnist_small.csv

\(~\)

Question #2 (Cross-validation)

Consider a toy data set of \(n=200\) observations generated via \(Y = f(X) + \epsilon\) where \(f(X) = 2x_1 + 10\) and \(\epsilon \sim N(0,\sigma =5)\). Notice that the true relationship between \(x_1\) and \(Y\) is linear, with 5 units of irreducible error. These data can be found at the URL given below:

https://remiller1450.github.io/data/toy_linear_data.csv

For the questions that follow you should use all \(n=200\) observations and do not need to perform a train-test split.

\(~\)

Question #3 (Application)

For this question you should use the dataset available here:

https://remiller1450.github.io/data/beans.csv

This data set was constructed using a computer vision system that segmented 13,611 images of 7 types of dry beans, extracting 16 features (12 dimensional features, 4 shape-form features). Additional details are contained in this paper

The following questions should be answered using Python code, including functions in the sklearn library. Unless otherwise indicated, use the 'accuracy' scoring criteria during model tuning and evaluation.

\(~\)