Course overview

Topic Math prerequisites Textbook references StatQuest videos
Linear regression Matrix notation, matrix–vector multiplication (IALA Sec. II.5)
Inner product/dot product (IALA Sec. I.1)
Norm of a vector (IALA Sec. I.3)
Matrix inverse (IALA Sec. II.11)
Derivatives and optimization (IALA App. C)
ISL: Sec. 3.1–3.3; Sec. 7.3 (basis functions)
ESL: Sec. 3.2 (linear regression and least squares); Sec. 5.1 (basis expansions)
PRML: Sec. 3.1, especially Sec. 3.1.1–3.1.2
IALA: Sec. 12.2 (OLS derivation)
Linear regression, clearly explained!!!
Multiple regression, clearly explained!!!
The Chain Rule
Gradient descent Derivatives and optimization (IALA App. C)
Complexity of algorithms, especially vector/matrix operations (IALA App. B, I.1, II.6)
ISL: Sec. 10.7.2 (SGD in a neural network context)
ESL: Sec. 10.10.1 (steepest descent in a gradient boosting context)
PRML: Sec. 3.1.3 (SGD and sequential learning for linear regression); Sec. 5.2.4 (general gradient-descent optimization)
Gradient descent, step-by-step
Error decomposition and bias–variance tradeoff Expectation of a random variable (linearity, constants)
Variance of a random variable
Independence of random variables
ISL: Sec. 2.2.2
ESL: Sec. 2.9; Sec. 7.2–7.3
PRML: Sec. 3.2
Machine Learning Fundamentals: Bias and Variance
Model selection ISL: Sec. 5.1.1–5.1.4 (cross-validation); Sec. 7.1, 7.3–7.4 (polynomial, basis, and spline model complexity)
ESL: Sec. 5.2 (piecewise polynomials and splines); Sec. 7.2–7.3, 7.10
PRML: Sec. 1.3
Cross validation
Regularization Derivatives and optimization (IALA App. C)
Norm of a vector (IALA Sec. I.3)
ISL: Sec. 6.2.1–6.2.3
ESL: Sec. 3.4.1–3.4.3
PRML: Sec. 3.1.4
Ridge regression (L2 regularization)
Lasso regression (L1 regularization)
Elasticnet regression (L1 and L2)
Logistic regression Probability mass functions
Bernoulli random variable
Independence of samples
Conditional probability
Joint probability
ISL: Sec. 4.3.1–4.3.5; Sec. 4.4.4 (Naive Bayes)
ESL: Sec. 4.4–4.4.1 (logistic regression); Sec. 6.6.3 (Naive Bayes)
PRML: Sec. 4.3.2–4.3.4; Sec. 4.2.3 (discrete generative classification and Naive-Bayes-style modeling)
Logistic Regression
Logistic Regression Details Pt1: Coefficients
Logistic Regression Details Pt 2: Maximum Likelihood
Odds and Log(Odds), Clearly Explained!!!
ROC and AUC, Clearly Explained!
Confusion matrix
Sensitivity and Specificity
K-nearest neighbor Expectation of a random variable (linearity, constants)
Variance of a random variable
Independence of random variables
ISL: Sec. 2.2.3, 3.5
ESL: Sec. 2.3.2; Sec. 13.3–13.5
PRML: Sec. 2.5.2 (nearest-neighbor methods)
K-nearest neighbors, Clearly Explained
Decision trees and ensembles Variance of a random variable
Independence of random variables
Variance of sum of random variables
ISL: Sec. 8.1.1–8.1.4, 8.2.1–8.2.3
ESL: Sec. 8.7 (bagging); Sec. 9.2 (trees); Sec. 10.1–10.12 (boosting); Sec. 15.1–15.4 (random forests)
PRML: Sec. 14.2–14.4 (committees, boosting, and tree-based models)
Decision Trees, Clearly Explained!!!
Decision Trees, Part 2 - Feature Selection and Missing Data
Regression Trees, Clearly Explained!!!
How to Prune Regression Trees, Clearly Explained!!!
Random Forests Part 1: Building, Using and Evaluating
AdaBoost, Clearly Explained
Support Vector Machines, Kernels, Other Kernel-Based Models Expectation of a random variable (linearity, constants)
Variance of a random variable
Independence of random variables
Conditional distributions
Gaussian distribution
ISL: Sec. 9.1–9.3
ESL: Sec. 12.2–12.3
PRML: Sec. 6.2 (kernels); Sec. 7.1–7.1.3 (maximum-margin/SVM classification); Sec. 6.4.2–6.4.3 (Gaussian-process regression and hyperparameter learning)
Support Vector Machines Part 1 (of 3): Main Ideas!!!
SVM with Polynomial kernel
SVM with RBF kernel
Neural networks Derivatives and optimization (IALA App. C)
Chain rule for multivariable functions
ISL: Sec. 10.1–10.2, 10.7.1
ESL: Sec. 11.3–11.5
PRML: Sec. 5.1–5.3
Neural Networks Pt. 1: Inside the Black Box
Neural Networks Pt. 2: Backpropagation Main Ideas
Backpropagation Details Part 1
Backpropagation Details Part 2
Neural Networks Pt. 3: ReLU In Action!!!
Neural Networks Pt. 4: Multiple Inputs and Outputs
Neural Networks Part 5: ArgMax and SoftMax
Neural Networks Part 6: Cross Entropy
Neural Networks Part 7: Cross Entropy Derivatives and Backpropagation
Introduction to PyTorch
Deep neural networks, convolutional neural networks ISL: Sec. 10.2, 10.3.1–10.3.5, 10.7.2–10.7.4, 10.8
ESL: Sec. 11.3–11.5 (classical neural network architecture and training only)
PRML: Sec. 5.5.6 (convolutional networks; older treatment)
Image Classification with Convolutional Neural Networks (CNNs)
Unsupervised learning Expectation and variance of random variables
Covariance and covariance matrices
Eigenvalues and eigenvectors
Probability distributions
Joint and conditional probability
ISL: Sec. 12.2.1–12.2.5 (PCA); Sec. 12.4.1–12.4.3 (K-means and hierarchical clustering)
ESL: Sec. 14.3.1–14.3.12 (clustering); Sec. 14.5.1–14.5.5 (PCA and related dimensionality reduction); Sec. 14.8–14.9 (further embeddings and dimensionality reduction)
PRML: Sec. 2.5.1, 9.2 (density estimation and Gaussian mixtures); Sec. 9.1 (K-means); Sec. 12.1.1–12.1.4 (PCA)
PCA main ideas in only 5 minutes!!!
Principal Component Analysis (PCA), Step-by-Step
K-means clustering
Word Embedding and Word2Vec, Clearly Explained!!!
Reinforcement learning Derivatives and optimization (IALA App. C)
Derivatives of exponential and log functions
Expectation of random variables
Conditional probability
RL: Sec. 2.1, 2.2, 2.4, 2.5, 2.8, 6.1, 6.5, 13.3 Reinforcement Learning: Essential Concepts
Recommender systems Matrix notation and matrix multiplication
Dot products and cosine similarity
Derivatives and optimization
Expectation of random variables
MMDS: Ch. 9, Recommendation Systems


Legend: