Assumes you know from ML-101
This lesson builds directly on ML-101 · Lesson 10 (Overfitting, Bias & Variance) and ML-101 · Lesson 4 (Logistic Regression). You should already be able to articulate the bias-variance decomposition of expected test error, and you should remember that logistic regression is fit by maximizing the log-likelihood with an L2 or L1 penalty added to the objective. We will re-derive those penalties from a Bayesian perspective, then turn to the more surprising phenomenon of double descent, where the classical U-shaped test-error curve is replaced by a non-monotone curve that has no optimum in the interior.
You should also remember from ML-101 · Lesson 11 (Gradient Descent & Optimization) that the optimizer does not know whether the L2 penalty it sees is "regularization" or part of the loss — both enter the gradient identically. That observation is exactly the lever for the AdamW derivation we will do here, but we will start from the prior-posterior derivation rather than from the optimizer.
Learning Objectives
- Derive L2 regularization as the maximum-a-posteriori (MAP) estimate under a Gaussian prior on the weights, and L1 as the MAP under a Laplace prior.
- Explain why L1 yields sparse solutions while L2 yields small but non-zero solutions, and connect this difference to the shape of the prior's level sets.
- Decompose the expected test error into bias, variance, and irreducible noise, and identify which term each regularizer affects.
- Describe the double-descent phenomenon and the assumption that the classical U-curve rests on; explain why modern overparameterized models break the assumption.
- Choose between L1, L2, early stopping, dropout, and weight decay for a given problem, and justify the choice in terms of the prior belief about the weights.
The Bayesian reframing: prior, likelihood, posterior
The classical supervised-learning setup is: given a dataset , find parameters that minimize a loss , typically the empirical risk
where is a per-example loss. The Bayesian reframing asks: what is the probability of given ? By Bayes' theorem,
The likelihood is the probability of the data under the model; if the labels are Gaussian with mean and variance , then . The prior is a belief about the weights before seeing the data. The posterior is the updated belief after seeing the data. The MAP estimate is
Substituting the Gaussian-likelihood expression,
This is the bridge between probabilistic and frequentist thinking: a prior on becomes a penalty term in the objective. Choosing the prior is choosing the regularizer.
L2 as a Gaussian prior
Suppose , a zero-mean isotropic Gaussian with covariance . Then
Substituting into the MAP objective,
The penalty coefficient is the ratio . It has a precise meaning: is how noisy the labels are believed to be; is how large the weights are believed to be a priori. A small (strong belief that weights are small) is a strong L2 penalty; a large (weak belief) is a weak penalty. This is the interpretation ML-101 omitted when it said "L2 prevents overfitting" — L2 encodes a specific prior, and the regularization strength is the prior precision.
The L2 solution has a closed form when (linear regression): the ridge-regression estimator
This is finite even when is singular — the term makes the system invertible. That is the algebraic reason L2 helps when features are collinear: it adds to every eigenvalue of , so the inverse exists.
L1 as a Laplace prior
Suppose instead independently for each coordinate — a zero-mean Laplace distribution. Then
The MAP objective becomes
L1 has no closed-form solution in general because is not differentiable at zero, but the geometry of L1 is what matters. The L1 unit ball has corners that touch the coordinate axes, while the L2 unit ball is round. When the unconstrained least-squares optimum lies near a coordinate axis, the L1-constrained solution lands exactly on the axis — i.e. that coordinate is zero. That is the source of L1 sparsity.
import numpy as np
def soft_threshold(x, lam):
"""The proximal operator for the L1 norm: argmin_z (1/2)(z - x)^2 + lam*|z|.
The closed-form solution sets z to zero when |x| <= lam (the L1 penalty
dominates and the unconstrained optimum is suppressed) and shrinks
|x| toward zero by exactly lam otherwise.
"""
return np.sign(x) * np.maximum(np.abs(x) - lam, 0.0)
# A coordinate descent step for Lasso: minimise (1/2n)||y - Xw||^2 + lam*||w||_1
# with respect to one coordinate w_j holding the others fixed.
def lasso_coordinate_step(X, y, w, j, lam):
r = y - X @ w + X[:, j] * w[j] # residual without feature j
rho = X[:, j] @ r # partial correlation of feature j with residual
w[j] = soft_threshold(rho / X[:, j] @ X[:, j], lam)
return w
The contrast between L1 and L2 is the cleanest demonstration of the geometric point above. L2 pulls every weight toward zero by a fixed proportion (and never quite reaches zero); L1 pulls every weight toward zero by a fixed amount, which means small weights are knocked all the way to zero and large weights are reduced by the same absolute amount.
| Penalty | Prior on | Behaviour at small | Behaviour at large | When to prefer |
|---|---|---|---|---|
| None | Flat (improper) | No shrinkage | No shrinkage | Plenty of data, strong prior belief that all features matter |
| L2 ($\lambda \ | w\ | _2^2$) | Gaussian | Small but non-zero |
| L1 ($\lambda \ | w\ | _1$) | Laplace | Sets exactly to zero |
| ElasticNet ($\alpha \ | w\ | _1 + (1-\alpha)\ | w\ | _2^2$) |
The bias-variance decomposition
The expected test error of an estimator fit on a training set of size drawn from a distribution is
The three terms are: bias-squared (how far the average prediction is from the truth), variance (how much a single fit's prediction moves around that average as the training set changes), and the irreducible noise in the labels.
Note what the expectation is over. The variance term is a variance of the estimator, not of the data: is the random training set, and the whole expression averages over both the fresh test point and the training draw. That is the difference between saying "the model has high variance" (a statement about one fit) and "the estimator has high variance" (a statement about the procedure). Only the second one decomposes this way.
This decomposition is what makes regularization sensible. A regularizer trades increased bias for reduced variance; the optimal strength is where the sum is maximised. The classical U-curve — test error falls as capacity increases, reaches a minimum, then rises as variance dominates — is exactly this bias-variance tradeoff drawn as a function of model capacity.
L2 reduces variance by shrinking the weights toward zero. It increases bias by pulling the average estimate away from the unconstrained least-squares solution. The two effects cross somewhere, and the test error is what tells you where. There is no closed-form way to find the crossing — that is why we tune regularization strength on a validation set.
Early stopping as implicit regularization
A different regularizer is to stop the optimizer before it has reached the minimum of the training loss. At step of gradient descent with learning rate , the parameters are approximately
where is the Hessian of the loss at the minimum . The matrix shrinks the initial residual along eigen-directions of , with effective shrinkage factor on the -th eigen-direction. Directions with large eigenvalues (steep curvature) shrink quickly; directions with small eigenvalues (flat curvature) shrink slowly and the iterate keeps moving along them.
This means early stopping implicitly performs coordinate-wise L2 shrinkage, with a different rate per eigen-direction than a uniform L2 penalty. It is, in a precise sense, an approximation to L2 with an eigenspectrum-dependent coefficient. This is also why "more data" and "earlier stopping" sometimes substitute for one another: both reduce the effective capacity of the model.
Double descent: the U-curve breaks
The classical U-curve assumes the model class is underparameterized relative to the data: there are more samples than effective parameters, and the model is identifiable. In that regime, increasing capacity strictly decreases bias and eventually increases variance, and the test error has a clean minimum. Modern deep networks violate this assumption. They are heavily overparameterized — a transformer with 7 billion parameters trained on 100 million tokens has many more parameters than samples — and the test error curve does not look like a U.
What happens instead is double descent. The test error first decreases as capacity grows (the classical regime), then increases as the model approaches the interpolation threshold (the point where the model can fit the training data perfectly), then decreases again as capacity grows further. The peak at the interpolation threshold is the "double-descent peak" and is the failure mode practitioners hit when their model is just large enough to memorize the training set but not yet large enough to find the smooth generalizing solution.
The mechanism is bias-variance breakdown near the interpolation threshold. Just below it, the model has to use all its parameters to fit the training set, so any small change in the data forces a large change in the parameters — variance is enormous. Just above it, there are many parameter configurations that fit the training set, and gradient descent finds one that lies in a low-variance basin of the loss landscape. The classical U-curve does not predict the second descent because it assumes a unique minimizer for each dataset, which is false in the overparameterized regime.
This is why the modern recipe is: pick a model much larger than the classical theory suggests, train it with strong regularization (weight decay, dropout, data augmentation), and stop when the validation loss stops improving. The model is in the second descent regime; the regularization keeps the iterate out of the high-variance basin near the interpolation threshold.
Dropout: a Bayesian approximation
Dropout trains an ensemble of networks in parallel: at each step, every neuron is independently zeroed out with probability . At test time, all neurons are used, but the weights are scaled by to keep the expected activation constant. The result is that the prediction is the geometric mean of an exponential number of "thinned" networks, which acts as a regularizer.
The Bayesian interpretation is that dropout is approximately variational inference in a Bernoulli belief network: each weight has a posterior distribution over "active" and "inactive", and dropout samples from that distribution at training time. The expected prediction under the dropout distribution is what the test-time scaling approximates.
Dropout is one form of regularization among many. For convolutional layers, it has largely been replaced by batch normalization (which has its own regularization effect through the noise in the minibatch statistics). For transformers, weight decay (AdamW) and stochastic depth (dropping entire layers) are the standard remedies. For tabular data, L2 and early stopping remain the defaults.
def dropout_forward(x, p, training=True):
"""Inverted dropout: scale by 1/(1-p) at training time, identity at test time."""
if not training:
return x
mask = (np.random.rand(*x.shape) >= p) / (1 - p)
return x * mask
The interaction between dropout and the optimizer is subtle. Dropout noise is not independent across steps — the same neuron can be active or inactive in a way that depends on the data and the previous activations — so it does not act like i.i.d. Gaussian noise on the gradient. The effective regularization strength is therefore not a simple function of ; you tune on a validation set, not by theory.
Elastic net: what the mixture actually buys you
Suppose two features are near-duplicates — say income and income_log — so their
design-matrix columns are almost collinear. Lasso's geometry then does something
unpleasant: the L1 ball touches one axis, the objective is nearly flat along the
direction that trades the two features off against each other, and coordinate descent
picks an essentially arbitrary one of them, dropping the other. The model is sparse, but
it is unstable sparse — re-run on a resampled dataset and a different feature survives.
The fix is not to abandon sparsity but to add a little L2 back:
The term makes the objective strictly convex and smooth, so the "nearly flat ridge" between correlated features is removed and the selected set becomes stable. then interpolates: is pure lasso, is pure ridge. What you have added is not a second regularizer for its own sake — it is a nonsmoothness remedy, and its price is that a genuinely useless feature can no longer be pushed to exactly zero.
The path is worth plotting. Fit the model over a grid of values and record which
features enter; features that are genuinely causal enter at every and form a
horizontal band, while features that enter only at near 1 are collinear
artifacts. In sklearn that is one line — ElasticNet(l1_ratio=...).path(x, y) — and it
is the closest thing to a stability-selection bootstrap that runs in seconds.
from sklearn.linear_model import ElasticNet
import numpy as np
alphas = np.logspace(-4, 0, 40) # sweep the overall strength
enet = ElasticNet(l1_ratio=0.5, max_iter=5000)
coefs = enet.path(X_train, y_train, coef_init=np.zeros(X_train.shape[1]),
alpha=alphas / alphas.max())[1]
# Columns that are non-zero for the whole sweep are the stable ones.
# Elastic net has no closed form, but it does have a coordinate-descent update that is
# the soft-threshold operator with a ridge correction:
def enet_coordinate_update(z, G_jj, alpha, l1_ratio, lam):
"""Minimise 0.5*G_jj*w_j^2 + l1_ratio*lam*|w_j| + 0.5*(1-l1_ratio)*lam*w_j^2
holding all other coordinates fixed. The penalty is quadratic-plus-L1, so the
solution is a soft-threshold scaled by the curvature the penalty itself adds.
"""
penalty = l1_ratio * lam
ridge = (1 - l1_ratio) * lam
# Effective curvature is the data curvature PLUS the ridge curvature.
return np.sign(z) * max(abs(z) - penalty, 0.0) / (G_jj + ridge)
Why L2 and AdamW are not the same thing
This is the single most common confusion carried over from ML-101, so it is worth being precise. With plain SGD, L2 regularization is implemented by adding the penalty to the gradient:
and the weight therefore decays at a rate proportional to its own magnitude — a multiplicative shrinkage. But add momentum or Adam and that equivalence breaks. Adam normalises the gradient by its second moment:
Now put the L2 term into . For a weight whose only gradient signal is the penalty, , so — the magnitude of cancels out completely. The update becomes a constant-size step toward zero, independent of how large the weight is. Coupling the penalty through Adam has therefore silently converted proportional shrinkage into sign-like shrinkage, which decays small weights far too slowly and changes the effective regularizer as training proceeds.
AdamW fixes this by moving the decay out of the gradient and into the parameter update:
The decay term now sits beside the Adam step rather than inside it, so it is not divided
by and retains its multiplicative character. The practical symptom of
getting this wrong: on a transformer fine-tune, L2-in-Adam under-regularises small
weights, the model fails to reach the sparsity you expected, and a sweep over weight_decay
produces a flat, confusing response.
Data augmentation as a prior over invariances
Augmentation is usually taught as a trick to manufacture more data. At 201 depth it is better read as an explicit statement of which transformations of the input you believe leave the label unchanged — a prior over invariances, in the same sense that L2 is a prior over weight magnitude.
Under the invariance hypothesis, if is a transformation with , then the correctly inflated training distribution is the one that marginalises over the group generated by . The model that minimises risk over that distribution is the one whose predictions are constant along group orbits. So augmentation is a way of encoding a symmetry into the hypothesis class rather than hoping SGD discovers it.
| Augmentation | Invariance asserted | Fails when |
|---|---|---|
| Horizontal flip | Camera orientation is irrelevant | The class encodes orientation (digits, traffic side) |
| Small translation | Object position is nuisance | Absolute position carries the label (geography, layout) |
| Colour jitter | Lighting is nuisance | Colour is the signal (disease grading, spectrograms) |
| CutMix / MixUp | Convexity of the data manifold | Labels are not linear in the inputs |
| Specular / time masks | Occlusion is transient | The occluded region is systematically informative |
That last column is the point. Every augmentation is a bet, and a bet is falsifiable. MixUp is not a free regularizer; its implicit assumption is that the data manifold is locally convex and that labels blend linearly. On a task where the classes are defined by a conjunction of features — "has both property A and property B" — a half-and-half image has no meaningful label and MixUp injects label noise. This is why augmentation ablations on a specific dataset are worth more than any prior belief about which one is better.
Label smoothing and what it really assumes
Label smoothing replaces the one-hot target with a mixture of the one-hot vector and the uniform distribution:
Minimising cross-entropy against pushes the model's predicted distribution toward the smoothed target rather than the exact one. The 101-level story is "it stops the model being overconfident". The precise story is what it does to the logit gap.
Let be the logit of the true class and the largest competitor. With no smoothing, the loss is unbounded below and the optimum is : the network is rewarded for driving the correct logit to ever-higher values even after the prediction is certainly right. Smoothing bounds it. Set at the smoothed optimum and you get
which is finite and, for and , is about . So label smoothing is a bounded-confidence prior: it caps how much probability mass the model may place on its most likely class, and the cap is a deterministic function of and . Two consequences that matter in practice. First, the mechanism is temperature-like, so it interacts with a temperature-scaled softmax at inference — you should not apply both blindly. Second, it is a prior on the output distribution, not a data augmentation, so it is entirely orthogonal to the input-side priors above; combining one does not substitute for the other.
Choosing a regularizer: a decision procedure
The order below is not arbitrary — it goes from assumptions you already believe, to knobs that cost the most to search.
- Start with none. If training error and test error are both high, the problem is underfitting and no regularizer will help. Adding one here makes it worse.
- Add early stopping. It is free, needs no new hyperparameter beyond the patience you already tuned, and it is the closest thing to a default.
- Add weight decay in its decoupled form (AdamW, or SGD with momentum where the equivalence does hold). This is the cheap, well-understood default for any model trained with an adaptive optimizer.
- Add data augmentation if you have a real invariance story. It is the only entry here that can improve the bias term rather than trading it away.
- Reach for L1/elastic net only when you specifically need variable selection on a tabular problem, and verify stability by re-running on resampled data.
- Use dropout for dense architectures where weight decay underperforms, and skip it entirely for transformers, where the standard stack is weight decay plus stochastic depth plus label smoothing.
def diagnose_bias_variance(X_tr, y_tr, X_va, y_va, factory, n_seeds=10):
"""
The only honest way to choose a regularizer: measure the two terms separately by
refitting on independent draws. A single train/validate split conflates them; this
averages the estimator's variance across resamples and exposes its bias as the gap
between the mean prediction and the validation truth.
"""
preds = np.array([factory().fit(X_tr, y_tr).predict(X_va) for _ in range(n_seeds)])
mean_pred = preds.mean(axis=0) # E_D[f(x; w_hat)] over resamples
bias_sq = ((mean_pred - y_va) ** 2).mean()
variance = preds.var(axis=0, ddof=1).mean() # Var_D[f(x; w_hat)]
noise = ((y_va - mean_pred) ** 2).mean() - variance # crude sigma^2 estimate
return {"bias^2": bias_sq, "variance": variance,
"irreducible": max(0.0, noise), "total": bias_sq + variance + max(0.0, noise)}
The resampling matters. Choosing a regularizer on a single validation split and then reporting that same split's error is the same mistake as tuning on the test set, and it is the reason "my validation curve looks U-shaped" is weak evidence for anything. The decomposition above needs independent draws to be meaningful at all.
Key Takeaways
- L2 regularization is the MAP estimate under a Gaussian prior on the weights; the regularization coefficient is the prior-precision ratio , not a free magic constant. L1 is the MAP under a Laplace prior and produces sparse solutions because the L1 ball has corners that touch coordinate axes.
- The bias-variance decomposition explains the classical U-curve: regularizers trade bias for variance, and the optimum is where the sum is maximised. This decomposition assumes an underparameterized model with a unique minimizer.
- Double descent is the failure of that assumption: in the overparameterized regime, the test-error curve is non-monotone, with a peak at the interpolation threshold. Modern deep learning lives in the second-descent regime; strong regularization keeps the iterate out of the high-variance basin.
- Early stopping is implicit per-eigendirection L2 shrinkage: directions of high curvature are shrunk fast, directions of low curvature are shrunk slowly. This is why early stopping is sometimes a substitute for explicit regularization.
- Dropout, weight decay, batch normalization, and stochastic depth are all regularizers with different assumptions. The choice depends on the prior belief about the model: dropout for overconfident deep nets, weight decay for transformers, batch normalization for convolutional architectures.