Course overview
| Topic | Math prerequisites | Textbook references | StatQuest videos |
|---|---|---|---|
| Linear regression |
Matrix notation, matrix–vector multiplication (IALA Sec. II.5) Inner product/dot product (IALA Sec. I.1) Norm of a vector (IALA Sec. I.3) Matrix inverse (IALA Sec. II.11) Derivatives and optimization (IALA App. C) |
ISL: Sec. 3.1–3.3; Sec. 7.3 (basis functions) ESL: Sec. 3.2 (linear regression and least squares); Sec. 5.1 (basis expansions) PRML: Sec. 3.1, especially Sec. 3.1.1–3.1.2 IALA: Sec. 12.2 (OLS derivation) |
Linear regression, clearly explained!!! Multiple regression, clearly explained!!! The Chain Rule |
| Gradient descent |
Derivatives and optimization (IALA App. C) Complexity of algorithms, especially vector/matrix operations (IALA App. B, I.1, II.6) |
ISL: Sec. 10.7.2 (SGD in a neural network context) ESL: Sec. 10.10.1 (steepest descent in a gradient boosting context) PRML: Sec. 3.1.3 (SGD and sequential learning for linear regression); Sec. 5.2.4 (general gradient-descent optimization) |
Gradient descent, step-by-step |
| Error decomposition and bias–variance tradeoff |
Expectation of a random variable (linearity, constants) Variance of a random variable Independence of random variables |
ISL: Sec. 2.2.2 ESL: Sec. 2.9; Sec. 7.2–7.3 PRML: Sec. 3.2 |
Machine Learning Fundamentals: Bias and Variance |
| Model selection | — |
ISL: Sec. 5.1.1–5.1.4 (cross-validation); Sec. 7.1, 7.3–7.4 (polynomial, basis, and spline model complexity) ESL: Sec. 5.2 (piecewise polynomials and splines); Sec. 7.2–7.3, 7.10 PRML: Sec. 1.3 |
Cross validation |
| Regularization |
Derivatives and optimization (IALA App. C) Norm of a vector (IALA Sec. I.3) |
ISL: Sec. 6.2.1–6.2.3 ESL: Sec. 3.4.1–3.4.3 PRML: Sec. 3.1.4 |
Ridge regression (L2 regularization) Lasso regression (L1 regularization) Elasticnet regression (L1 and L2) |
| Logistic regression |
Probability mass functions Bernoulli random variable Independence of samples Conditional probability Joint probability |
ISL: Sec. 4.3.1–4.3.5; Sec. 4.4.4 (Naive Bayes) ESL: Sec. 4.4–4.4.1 (logistic regression); Sec. 6.6.3 (Naive Bayes) PRML: Sec. 4.3.2–4.3.4; Sec. 4.2.3 (discrete generative classification and Naive-Bayes-style modeling) |
Logistic Regression Logistic Regression Details Pt1: Coefficients Logistic Regression Details Pt 2: Maximum Likelihood Odds and Log(Odds), Clearly Explained!!! ROC and AUC, Clearly Explained! Confusion matrix Sensitivity and Specificity |
| K-nearest neighbor |
Expectation of a random variable (linearity, constants) Variance of a random variable Independence of random variables |
ISL: Sec. 2.2.3, 3.5 ESL: Sec. 2.3.2; Sec. 13.3–13.5 PRML: Sec. 2.5.2 (nearest-neighbor methods) |
K-nearest neighbors, Clearly Explained |
| Decision trees and ensembles |
Variance of a random variable Independence of random variables Variance of sum of random variables |
ISL: Sec. 8.1.1–8.1.4, 8.2.1–8.2.3 ESL: Sec. 8.7 (bagging); Sec. 9.2 (trees); Sec. 10.1–10.12 (boosting); Sec. 15.1–15.4 (random forests) PRML: Sec. 14.2–14.4 (committees, boosting, and tree-based models) |
Decision Trees, Clearly Explained!!! Decision Trees, Part 2 - Feature Selection and Missing Data Regression Trees, Clearly Explained!!! How to Prune Regression Trees, Clearly Explained!!! Random Forests Part 1: Building, Using and Evaluating AdaBoost, Clearly Explained |
| Support Vector Machines, Kernels, Other Kernel-Based Models |
Expectation of a random variable (linearity, constants) Variance of a random variable Independence of random variables Conditional distributions Gaussian distribution |
ISL: Sec. 9.1–9.3 ESL: Sec. 12.2–12.3 PRML: Sec. 6.2 (kernels); Sec. 7.1–7.1.3 (maximum-margin/SVM classification); Sec. 6.4.2–6.4.3 (Gaussian-process regression and hyperparameter learning) |
Support Vector Machines Part 1 (of 3): Main Ideas!!! SVM with Polynomial kernel SVM with RBF kernel |
| Neural networks |
Derivatives and optimization (IALA App. C) Chain rule for multivariable functions |
ISL: Sec. 10.1–10.2, 10.7.1 ESL: Sec. 11.3–11.5 PRML: Sec. 5.1–5.3 |
Neural Networks Pt. 1: Inside the Black Box Neural Networks Pt. 2: Backpropagation Main Ideas Backpropagation Details Part 1 Backpropagation Details Part 2 Neural Networks Pt. 3: ReLU In Action!!! Neural Networks Pt. 4: Multiple Inputs and Outputs Neural Networks Part 5: ArgMax and SoftMax Neural Networks Part 6: Cross Entropy Neural Networks Part 7: Cross Entropy Derivatives and Backpropagation Introduction to PyTorch |
| Deep neural networks, convolutional neural networks | — |
ISL: Sec. 10.2, 10.3.1–10.3.5, 10.7.2–10.7.4, 10.8 ESL: Sec. 11.3–11.5 (classical neural network architecture and training only) PRML: Sec. 5.5.6 (convolutional networks; older treatment) |
Image Classification with Convolutional Neural Networks (CNNs) |
| Unsupervised learning |
Expectation and variance of random variables Covariance and covariance matrices Eigenvalues and eigenvectors Probability distributions Joint and conditional probability |
ISL: Sec. 12.2.1–12.2.5 (PCA); Sec. 12.4.1–12.4.3 (K-means and hierarchical clustering) ESL: Sec. 14.3.1–14.3.12 (clustering); Sec. 14.5.1–14.5.5 (PCA and related dimensionality reduction); Sec. 14.8–14.9 (further embeddings and dimensionality reduction) PRML: Sec. 2.5.1, 9.2 (density estimation and Gaussian mixtures); Sec. 9.1 (K-means); Sec. 12.1.1–12.1.4 (PCA) |
PCA main ideas in only 5 minutes!!! Principal Component Analysis (PCA), Step-by-Step K-means clustering Word Embedding and Word2Vec, Clearly Explained!!! |
| Reinforcement learning |
Derivatives and optimization (IALA App. C) Derivatives of exponential and log functions Expectation of random variables Conditional probability |
RL: Sec. 2.1, 2.2, 2.4, 2.5, 2.8, 6.1, 6.5, 13.3 |
Reinforcement Learning: Essential Concepts |
| Recommender systems |
Matrix notation and matrix multiplication Dot products and cosine similarity Derivatives and optimization Expectation of random variables |
MMDS: Ch. 9, Recommendation Systems | — |
Legend:
- ISL – Introduction to Statistical Learning (James et al.)
- ESL – Elements of Statistical Learning (Hastie, Tibshirani, Friedman)
- PRML – Pattern Recognition and Machine Learning (Bishop)
- IALA – Introduction to Applied Linear Algebra (Boyd & Vandenberghe)
- MMDS – Mining of Massive Datasets (Leskovec, Rajaraman, Ullman)
- RL – Reinforcement Learning: An Introduction (Sutton & Barto)