Assumes you know from ML-101
You have completed ML-101. Specifically, you have met gradient descent in ML-101 lesson 11 — Gradient Descent & Optimization, the bias-variance picture in ML-101 lesson 10 — Overfitting, Bias & Variance, the basics of cross-entropy and regularization in ML-101 lesson 4 — Logistic Regression and ML-101 lesson 14 — Neural Networks & Backprop, and the deployment mindset of ML-101 lesson 16 — MLOps & Deployment. This lesson does not re-derive those facts from scratch. Instead it restates the one thing you will build on — that empirical risk minimization is a discrete optimization over a hypothesis class parameterized by real numbers — and then argues that deriving every subsequent rule from that one fact, rather than asserting it, is the difference between ML-101 and ML-201.
Learning Objectives
- State, in one sentence, the single optimisation problem that every supervised model in ML-101 solves, and identify the three places where that statement silently hides a choice (the loss, the parameterisation, the regularizer).
- For any rule from ML-101 — L2 regularization, the learning-rate schedule, the softmax link, early stopping, min-max normalisation — name the underlying assumption whose violation breaks the rule.
- Explain why "works on the validation set" is an unreliable statement once the model itself was tuned against the validation set, and what the fix is.
- Convert a 101-level claim such as "cross-entropy is better than MSE for classification" into a 201-level derivation that names the noise model implied by each loss.
- Read a researcher's footnote about a result holding "in expectation" and translate that into the finite-sample statement they actually proved.
The single fact everything else stands on
Every model you met in ML-101 — logistic regression, decision trees, k-NN, the small neural network — can be written as the same five-symbol optimisation problem:
The four moving parts are: a hypothesis class , a loss , a regularizer and its weight , and an optimiser that does not appear in the formula but is implicit in how you intend to find . None of those four parts is given by the data. The data only specifies that something in must be picked; everything else is a choice. The whole of ML-201 is the study of those choices, viewed from the question "which assumption about the world makes this choice correct?".
A reader who treats the formula as a black box is doing ML-101. A reader who can hear every silence in it — what kind of loss is allowed here, what regularizer is compatible with this loss, what optimiser converges on this regularizer — is doing ML-201. The gap is not depth of knowledge; it is the habit of noticing the choices.
Why ML-201 derives and ML-101 asserts
ML-101 is a vocabulary course. It teaches the names: gradient, cross-entropy, regularization, overfitting, convergence. Naming is necessary because every later concept is built from those names. ML-201 is a forensic course: it asks "given that the rule holds, what would the world have to look like?" and then "given that the world is the way it actually is, in what regime does the rule fail?".
The standard ML-101 footnote — "the test set should be drawn from the same distribution as the training set" — is an assertion in ML-101 and a derivation in ML-201. In ML-201 the assertion is re-stated as a hypothesis test of against the alternative that they differ on some measurable set :
If the empirical estimate of the squared maximum mean discrepancy exceeds the quantile of its null distribution, the i.i.d. assumption underlying every generalisation bound in ML-101 is rejected. The same data, the same model, the same validation accuracy — but the conclusion inverts.
This is not a curiosity. It is the work ML-201 exists to do. Every rule from ML-101 has a latent distributional assumption; the entire course is the work of naming those assumptions and showing what happens at the boundary.
The three hidden choices
Look again at the five-symbol optimisation. Three choices are silently bundled in.
Choice 1 — the loss is a likelihood in disguise. Almost every loss used in ML-101 — squared error, cross-entropy, hinge, the Poisson log-loss — is the negative log-likelihood of some conditional model . The mapping is exact, not approximate: squared error implies Gaussian noise of constant variance, cross-entropy implies a categorical/Cox–Huels distribution, hinge implies a label noise of , Poisson loss implies count noise. If you change the loss without changing the assumed noise, the new loss is just an arbitrary penalty and the model has lost its probabilistic meaning.
Choice 2 — the parameterisation is the hypothesis class. A "linear model" written as and the same model written as with the constraint and the additional constraint are the same function on the data and a different optimisation problem for the gradient descent routine. The choice of parameterisation changes the optimiser's path, the conditioning of the Hessian, and the shape of the basin. ML-101 said "use a small model if you don't have much data". ML-201 says: the same model class, written in two different parameterisations, will behave like two different models under finite-step optimisation.
Choice 3 — the regularizer encodes a prior. is rarely "there to prevent overfitting" in the textbook sense. It is the negative log of a prior distribution on parameters: L2 corresponds to a Gaussian prior and L1 to a Laplace prior . The posterior mode under those priors is the L2- or L1-penalised empirical risk minimiser. Choosing L2 is choosing the prior that parameter values cluster near zero. Choosing no regulariser is choosing the improper uniform prior, which has measure-theoretic consequences that bite you in Bayesian model comparison.
| Choice | What ML-101 calls it | What ML-201 says it is |
|---|---|---|
| "Loss function" | Negative log-likelihood of a noise model | |
| "Regularizer" | Negative log of a prior on | |
| "Regularization strength" | The signal-to-noise ratio of the prior to the likelihood |
Why the bias-variance picture is not symmetric
The classical ML-101 lesson 10 picture of bias and variance comes from a fixed-design regression setup, with and . There, the expected prediction error at decomposes as
The ML-101 reading is "we can trade bias for variance". The 201 reading is that this decomposition assumes the model class is rich enough that the bias term can be driven toward zero (otherwise the trade-off is not a trade-off but a floor), and assumes the variance term is finite (otherwise the prediction has no expectation at all and the equation is not a decomposition — it is a divergence). The picture breaks twice:
- Over-parameterised models. A neural network with more parameters than data points has zero training error, near-zero bias on the training distribution, and a non-trivial variance. The decomposition still holds in the limit of infinite model width, but in the finite-width, finite-data regime the variance term depends on the path of the optimiser, not just on the model class.
- Quantile losses and asymmetric costs. The mean-squared-error decomposition exploits the fact that the conditional mean is the Bayes-optimal under squared loss. Under quantile or asymmetric loss the optimal predictor is the conditional quantile, and the decomposition becomes conditional-quantile / quantile-crossing / variance — a different picture, not a re-labelling of the same one.
These failure modes are not esoteric. They appear in any real ML-201 problem the moment the model is bigger than the data or the cost is asymmetric. Recognising them requires nothing more than re-reading the decomposition and asking: "which line of this equation requires the assumption I'm trying to relax?"
What "in expectation" really means
Many results in ML-101 — the generalisation bound, the bias-variance decomposition, the central-limit behaviour of the SGD iterate — are stated in expectation. A careful reader notices that the finite-sample statement is almost always weaker than the asymptotic one. The rule of thumb is that a result proved in expectation over the data-generating distribution gives a statement about a hypothetical average run, not about the run you have on disk.
A practitioner who treats "in expectation" as "for me" is the practitioner who reports a 95 % confidence interval and watches the next experiment land three standard deviations away. The conversion is mechanical: if a result states
then the finite-sample tail statement, given a concentration inequality (McDiarmid, Bernstein, or the bounded-differences inequality depending on which gradient is bounded), looks like
where is the bound on the per-example loss difference when one example is replaced. The "in expectation" claim and the "with high probability" claim are not the same statement, and one does not imply the other without paying the cost in the exponent. This is the difference between a 101-level "expected test error 5 %" and a 201-level "with at least 95 % probability over realisations of , the test error of is at most 9 % on examples drawn from a -bounded-loss distribution".
The McDiarmid-style concentration argument is constructive; you can make it concrete in a few lines.
import math
def finite_sample_bound(n: int, c: float, delta: float = 0.05) -> float:
"""Slack t such that R(f) <= E[R(f)] + t with probability >= 1 - delta.
McDiarmid's inequality says P[|f - E[f]| >= t] <= 2 exp(-2 n t^2 / c^2),
so for a one-sided (delta) bound on the upper tail we need
2 exp(-2 n t^2 / c^2) == delta
=> t == c * sqrt(log(2/delta) / (2 n)).
"""
return c * math.sqrt(math.log(2.0 / delta) / (2.0 * n))
# With c=0.1 (bounded per-example loss differences) and n=10_000,
# the slack is about 0.014 — small, but multiplicative across many models.
for n in [1_000, 10_000, 100_000]:
print(n, round(finite_sample_bound(n, c=0.1), 5))
The 201 lesson is to report this slack, not to omit it.
The validation set is a consumable
The single most expensive 101-level mistake is to read the validation accuracy as a property of the model. The validation accuracy is a property of the triple (model, hyperparameters, validation set). Once you have used the validation set times to pick candidate hyperparameters, you have implicitly selected from a search space of size where is the per-trial hyperparameter grid. The empirical validation accuracy is the maximum of i.i.d. random variables, each identically distributed to the true generalisation accuracy.
The expected maximum of iid draws of a -valued quantity with CDF is
For candidates with the true generalisation accuracy at , the expected maximum is already . After candidates it is . The validation accuracy must over-report generalisation once you tune against it, and the magnitude of the bias is a function of the search budget, not of the model. The fix is nested cross-validation or a held-out test set that has been touched exactly once.
This is the topic of ML-201 lesson 8 in detail, but the orientation is given here because every other lesson in the course silently assumes you already believe it.
A sanity check: predicting the expected maximum of iid draws
The expected-maximum calculation in the previous section is not folklore; it is a direct consequence of order statistics and you can compute it in three lines.
import numpy as np
from scipy.stats import beta
def expected_max(k: int, true_acc: float, n_mc: int = 200_000) -> float:
"""Monte-Carlo estimate of E[max(M_1, ..., M_k)] for M_i ~ Bernoulli(true_acc)."""
# Each M_i is 1 with probability true_acc; the empirical maximum is
# the fraction of k-draw batches where at least one draw hits the rare class.
samples = np.random.binomial(1, true_acc, size=(n_mc, k)).max(axis=1)
return float(samples.mean())
# As k grows, expected_max climbs even though the underlying accuracy is fixed.
for k in [1, 5, 20, 100, 500]:
print(k, round(expected_max(k, 0.80), 4))
Running this with true_acc = 0.80 reproduces the table in the previous
section: at trials the empirical maximum validation accuracy is
already , and at it is . The point is not
the specific numbers — it is that the bias is a function of the search budget,
not of the model. The same model, tuned on candidates, reports ;
tuned on candidates, it reports . The validation
accuracy is not the model's property; it is the search's property.
A worked re-derivation: why cross-entropy, not MSE
ML-101 lesson 4 — Logistic Regression said: "use cross-entropy for classification". ML-201 derives it.
Suppose the labels are draws from a categorical distribution with class probabilities . The likelihood of the dataset is
The negative log-likelihood is
which is exactly cross-entropy. Now replace the noise model. Suppose instead that the labels are real-valued and Gaussian around a sigmoid:
Then the negative log-likelihood is
which is squared error. Both losses are maximum likelihood; they differ only in the assumed noise. The 101-level claim "MSE is wrong for classification" is a 201-level claim that the Gaussian noise model is wrong for binary labels — which is true, because the support of a Gaussian is and the support of a binary label is .
import numpy as np
def sigmoid(z: np.ndarray) -> np.ndarray:
# Numerically stable sigmoid; the naive 1/(1+exp(-z)) overflows for z << 0.
out = np.empty_like(z)
pos = z >= 0
out[pos] = 1.0 / (1.0 + np.exp(-z[pos]))
ex = np.exp(z[~pos])
out[~pos] = ex / (1.0 + ex)
return out
def nll_cross_entropy(y: np.ndarray, p: np.ndarray) -> float:
# Categorical noise model: p is the predicted P(y=1|x).
p = np.clip(p, 1e-12, 1 - 1e-12) # log(0) is -inf; clip keeps the loss finite.
return float(-(y * np.log(p) + (1 - y) * np.log(1 - p)).mean())
def nll_squared_error(y: np.ndarray, p: np.ndarray, sigma: float = 1.0) -> float:
# Gaussian noise model around the same predicted probability.
return float(((y - p) ** 2).mean() / (2 * sigma ** 2))
# Toy demo: same predictions, two losses. The losses are NOT comparable in magnitude,
# because they live on different log-likelihood scales; only their GRADIENTS share units.
rng = np.random.default_rng(0)
y = rng.integers(0, 2, size=1000).astype(float)
p = sigmoid(2 * rng.standard_normal(1000))
print("cross-entropy nll:", round(nll_cross_entropy(y, p), 4))
print("squared-error nll:", round(nll_squared_error(y, p), 4))
The 201-level follow-up is more interesting: if you know the noise, the loss is determined; if you do not, then the "loss function" is just a penalty and you cannot interpret the value of the loss as a likelihood. This is why calibration metrics exist, and it is why the test of a good loss is whether the predictions it produces are well-calibrated probabilities, not whether the loss goes down.
Key Takeaways
- The single optimisation problem at the back of ML-101 hides three choices (loss, parameterisation, regularizer); each is a probabilistic statement about the world, not a free knob.
- "In expectation" is not "for me"; the conversion to a finite-sample statement costs in the exponent.
- The validation set is a consumable — every touch of it inflates the expected maximum of the validation accuracy by a computable amount.
- Cross-entropy is not "punishing" — it is the negative log-likelihood under categorical noise. MSE is the negative log-likelihood under Gaussian noise. The loss encodes the noise model.
- ML-201 is the discipline of reading the silence in every 101-level assertion: which assumption, if violated, breaks the result.