06

Regularization & Generalization

Cantonese podcast title: 正則化與泛化

Learning Objectives

  1. Derive L2 regularization as the maximum-a-posteriori (MAP) estimate under a Gaussian prior on the weights, and L1 as the MAP under a Laplace prior.
  2. Explain why L1 yields sparse solutions while L2 yields small but non-zero solutions, and connect this difference to the shape of the prior's level sets.
  3. Decompose the expected test error into bias, variance, and irreducible noise, and identify which term each regularizer affects.
  4. Describe the double-descent phenomenon and the assumption that the classical U-curve rests on; explain why modern overparameterized models break the assumption.
  5. Choose between L1, L2, early stopping, dropout, and weight decay for a given problem, and justify the choice in terms of the prior belief about the weights.
Regularization & Generalization — visual guide
Double descent: classical U-curve and modern non-monotone test error Test error vs model capacity: classical U vs double descent the interpolation threshold (n_train samples) is where the U-curve breaks classical U-curve underparameterized regime double descent interpolation threshold = n_train samples peak: many high-variance zero-loss solutions exist model capacity (parameters / samples) test error Regularizer cheat sheet L2 (Gaussian prior) ||w||_2^2 small, never zero use for: smooth weights L1 (Laplace prior) sparse, exact zeros use for: many features Early stopping per-eigendirection L2 use for: cheap, no prior model Dropout / weight decay keep iterate in low-variance basin in second-descent regime Modern deep learning lives in the second descent. Strong regularization keeps the iterate out of the high-variance basin near the interpolation threshold.

Assumes you know from ML-101

This lesson builds directly on ML-101 · Lesson 10 (Overfitting, Bias & Variance) and ML-101 · Lesson 4 (Logistic Regression). You should already be able to articulate the bias-variance decomposition of expected test error, and you should remember that logistic regression is fit by maximizing the log-likelihood with an L2 or L1 penalty added to the objective. We will re-derive those penalties from a Bayesian perspective, then turn to the more surprising phenomenon of double descent, where the classical U-shaped test-error curve is replaced by a non-monotone curve that has no optimum in the interior.

You should also remember from ML-101 · Lesson 11 (Gradient Descent & Optimization) that the optimizer does not know whether the L2 penalty it sees is "regularization" or part of the loss — both enter the gradient identically. That observation is exactly the lever for the AdamW derivation we will do here, but we will start from the prior-posterior derivation rather than from the optimizer.

Learning Objectives

  1. Derive L2 regularization as the maximum-a-posteriori (MAP) estimate under a Gaussian prior on the weights, and L1 as the MAP under a Laplace prior.
  2. Explain why L1 yields sparse solutions while L2 yields small but non-zero solutions, and connect this difference to the shape of the prior's level sets.
  3. Decompose the expected test error into bias, variance, and irreducible noise, and identify which term each regularizer affects.
  4. Describe the double-descent phenomenon and the assumption that the classical U-curve rests on; explain why modern overparameterized models break the assumption.
  5. Choose between L1, L2, early stopping, dropout, and weight decay for a given problem, and justify the choice in terms of the prior belief about the weights.

The Bayesian reframing: prior, likelihood, posterior

The classical supervised-learning setup is: given a dataset D={(xi,yi)}i=1n\mathcal{D} = \{(x_i, y_i)\}_{i=1}^n, find parameters ww that minimize a loss L(w;D)L(w; \mathcal{D}), typically the empirical risk

L(w)=1n∑i=1nℓ(yi,f(xi;w))L(w) = \frac{1}{n} \sum_{i=1}^n \ell(y_i, f(x_i; w))

where ℓ\ell is a per-example loss. The Bayesian reframing asks: what is the probability of ww given D\mathcal{D}? By Bayes' theorem,

p(w∣D)  =  p(D∣w) p(w)p(D)p(w \mid \mathcal{D}) \;=\; \frac{p(\mathcal{D} \mid w)\, p(w)}{p(\mathcal{D})}

The likelihood p(D∣w)p(\mathcal{D} \mid w) is the probability of the data under the model; if the labels are Gaussian with mean f(xi;w)f(x_i; w) and variance σ2\sigma^2, then p(D∣w)∝exp⁡(−12σ2∑i(yi−f(xi;w))2)p(\mathcal{D} \mid w) \propto \exp\bigl(-\tfrac{1}{2\sigma^2} \sum_i (y_i - f(x_i; w))^2\bigr). The prior p(w)p(w) is a belief about the weights before seeing the data. The posterior p(w∣D)p(w \mid \mathcal{D}) is the updated belief after seeing the data. The MAP estimate is

w^MAP  =  arg⁡max⁡w  log⁡p(w∣D)  =  arg⁡max⁡w  [log⁡p(D∣w)+log⁡p(w)]\hat{w}_{\text{MAP}} \;=\; \arg\max_w \; \log p(w \mid \mathcal{D}) \;=\; \arg\max_w \;\bigl[ \log p(\mathcal{D} \mid w) + \log p(w) \bigr]

Substituting the Gaussian-likelihood expression,

w^MAP  =  arg⁡min⁡w  [12σ2∑i(yi−f(xi;w))2  −  log⁡p(w)]\hat{w}_{\text{MAP}} \;=\; \arg\min_w \;\Bigl[ \frac{1}{2\sigma^2} \sum_i (y_i - f(x_i; w))^2 \;-\; \log p(w) \Bigr]

This is the bridge between probabilistic and frequentist thinking: a prior on ww becomes a penalty term in the objective. Choosing the prior is choosing the regularizer.

L2 as a Gaussian prior

Suppose p(w)=N(0,τ2I)p(w) = \mathcal{N}(0, \tau^2 I), a zero-mean isotropic Gaussian with covariance τ2I\tau^2 I. Then

log⁡p(w)  =  −12τ2∥w∥22  +  const.\log p(w) \;=\; -\frac{1}{2\tau^2} \|w\|_2^2 \;+\; \text{const.}

Substituting into the MAP objective,

w^MAP  =  arg⁡min⁡w  [∑i(yi−f(xi;w))2  +  σ2τ2∥w∥22]\hat{w}_{\text{MAP}} \;=\; \arg\min_w \;\Bigl[ \sum_i (y_i - f(x_i; w))^2 \;+\; \frac{\sigma^2}{\tau^2} \|w\|_2^2 \Bigr]

The penalty coefficient is the ratio σ2/τ2\sigma^2 / \tau^2. It has a precise meaning: σ2\sigma^2 is how noisy the labels are believed to be; τ2\tau^2 is how large the weights are believed to be a priori. A small τ2\tau^2 (strong belief that weights are small) is a strong L2 penalty; a large τ2\tau^2 (weak belief) is a weak penalty. This is the interpretation ML-101 omitted when it said "L2 prevents overfitting" — L2 encodes a specific prior, and the regularization strength is the prior precision.

The L2 solution has a closed form when f(x;w)=w⊤xf(x; w) = w^\top x (linear regression): the ridge-regression estimator

w^ridge  =  (X⊤X+λI)−1X⊤y,λ=σ2τ2\hat{w}_{\text{ridge}} \;=\; (X^\top X + \lambda I)^{-1} X^\top y, \quad \lambda = \frac{\sigma^2}{\tau^2}

This is finite even when X⊤XX^\top X is singular — the λI\lambda I term makes the system invertible. That is the algebraic reason L2 helps when features are collinear: it adds λ\lambda to every eigenvalue of X⊤XX^\top X, so the inverse exists.

L1 as a Laplace prior

Suppose instead p(w)=12bexp⁡(−∣wi∣/b)p(w) = \frac{1}{2b} \exp\bigl(-\lvert w_i \rvert / b\bigr) independently for each coordinate — a zero-mean Laplace distribution. Then

log⁡p(w)  =  −1b∑i∣wi∣  +  const.  =  −1b∥w∥1  +  const.\log p(w) \;=\; -\frac{1}{b} \sum_i \lvert w_i \rvert \;+\; \text{const.} \;=\; -\frac{1}{b} \|w\|_1 \;+\; \text{const.}

The MAP objective becomes

w^MAP  =  arg⁡min⁡w  [∑i(yi−f(xi;w))2  +  σ2b∥w∥1]\hat{w}_{\text{MAP}} \;=\; \arg\min_w \;\Bigl[ \sum_i (y_i - f(x_i; w))^2 \;+\; \frac{\sigma^2}{b} \|w\|_1 \Bigr]

L1 has no closed-form solution in general because ∣w∣\lvert w \rvert is not differentiable at zero, but the geometry of L1 is what matters. The L1 unit ball {w:∥w∥1≤c}\lbrace w : \|w\|_1 \le c \rbrace has corners that touch the coordinate axes, while the L2 unit ball is round. When the unconstrained least-squares optimum lies near a coordinate axis, the L1-constrained solution lands exactly on the axis — i.e. that coordinate is zero. That is the source of L1 sparsity.

import numpy as np

def soft_threshold(x, lam):
    """The proximal operator for the L1 norm: argmin_z (1/2)(z - x)^2 + lam*|z|.

    The closed-form solution sets z to zero when |x| <= lam (the L1 penalty
    dominates and the unconstrained optimum is suppressed) and shrinks
    |x| toward zero by exactly lam otherwise.
    """
    return np.sign(x) * np.maximum(np.abs(x) - lam, 0.0)

# A coordinate descent step for Lasso: minimise (1/2n)||y - Xw||^2 + lam*||w||_1
# with respect to one coordinate w_j holding the others fixed.
def lasso_coordinate_step(X, y, w, j, lam):
    r = y - X @ w + X[:, j] * w[j]       # residual without feature j
    rho = X[:, j] @ r                     # partial correlation of feature j with residual
    w[j] = soft_threshold(rho / X[:, j] @ X[:, j], lam)
    return w

The contrast between L1 and L2 is the cleanest demonstration of the geometric point above. L2 pulls every weight toward zero by a fixed proportion (and never quite reaches zero); L1 pulls every weight toward zero by a fixed amount, which means small weights are knocked all the way to zero and large weights are reduced by the same absolute amount.

PenaltyPrior on wiw_iBehaviour at small ∣wi∣\lvert w_i \rvertBehaviour at large ∣wi∣\lvert w_i \rvertWhen to prefer
NoneFlat (improper)No shrinkageNo shrinkagePlenty of data, strong prior belief that all features matter
L2 ($\lambda \w\_2^2$)Gaussian N(0,τ2)\mathcal{N}(0, \tau^2)Small but non-zero
L1 ($\lambda \w\_1$)Laplace(0,b)(0, b)Sets exactly to zero
ElasticNet ($\alpha \w\_1 + (1-\alpha)\w\_2^2$)

The bias-variance decomposition

The expected test error of an estimator w^\hat{w} fit on a training set of size nn drawn from a distribution PP is

E(x,y)∼P, D[(y−f(x;w^))2]  =  (f(x)−ED[f(x;w^)])2⏟bias2  +  Var⁡D[f(x;w^)]⏟variance  +  σ2\mathbb{E}_{(x,y) \sim P,\, \mathcal{D}} \bigl[ (y - f(x; \hat{w}))^2 \bigr] \;=\; \underbrace{\bigl( f(x) - \mathbb{E}_\mathcal{D}[f(x; \hat{w})] \bigr)^2}_{\text{bias}^2} \;+\; \underbrace{\operatorname{Var}_\mathcal{D}\bigl[f(x; \hat{w})\bigr]}_{\text{variance}} \;+\; \sigma^2

The three terms are: bias-squared (how far the average prediction is from the truth), variance (how much a single fit's prediction moves around that average as the training set changes), and the irreducible noise σ2\sigma^2 in the labels.

Note what the expectation is over. The variance term is a variance of the estimator, not of the data: D\mathcal{D} is the random training set, and the whole expression averages over both the fresh test point and the training draw. That is the difference between saying "the model has high variance" (a statement about one fit) and "the estimator has high variance" (a statement about the procedure). Only the second one decomposes this way.

This decomposition is what makes regularization sensible. A regularizer trades increased bias for reduced variance; the optimal strength is where the sum is maximised. The classical U-curve — test error falls as capacity increases, reaches a minimum, then rises as variance dominates — is exactly this bias-variance tradeoff drawn as a function of model capacity.

L2 reduces variance by shrinking the weights toward zero. It increases bias by pulling the average estimate away from the unconstrained least-squares solution. The two effects cross somewhere, and the test error is what tells you where. There is no closed-form way to find the crossing — that is why we tune regularization strength on a validation set.

Early stopping as implicit regularization

A different regularizer is to stop the optimizer before it has reached the minimum of the training loss. At step tt of gradient descent with learning rate η\eta, the parameters are approximately

wt  ≈  w∞  −  (I−ηH)t(w∞−w0)w_t \;\approx\; w_\infty \;-\; (I - \eta H)^t (w_\infty - w_0)

where HH is the Hessian of the loss at the minimum w∞w_\infty. The matrix (I−ηH)t(I - \eta H)^t shrinks the initial residual (w∞−w0)(w_\infty - w_0) along eigen-directions of HH, with effective shrinkage factor (1−ηλi)t(1 - \eta \lambda_i)^t on the ii-th eigen-direction. Directions with large eigenvalues (steep curvature) shrink quickly; directions with small eigenvalues (flat curvature) shrink slowly and the iterate keeps moving along them.

This means early stopping implicitly performs coordinate-wise L2 shrinkage, with a different rate per eigen-direction than a uniform L2 penalty. It is, in a precise sense, an approximation to L2 with an eigenspectrum-dependent coefficient. This is also why "more data" and "earlier stopping" sometimes substitute for one another: both reduce the effective capacity of the model.

Double descent: the U-curve breaks

The classical U-curve assumes the model class is underparameterized relative to the data: there are more samples than effective parameters, and the model is identifiable. In that regime, increasing capacity strictly decreases bias and eventually increases variance, and the test error has a clean minimum. Modern deep networks violate this assumption. They are heavily overparameterized — a transformer with 7 billion parameters trained on 100 million tokens has many more parameters than samples — and the test error curve does not look like a U.

What happens instead is double descent. The test error first decreases as capacity grows (the classical regime), then increases as the model approaches the interpolation threshold (the point where the model can fit the training data perfectly), then decreases again as capacity grows further. The peak at the interpolation threshold is the "double-descent peak" and is the failure mode practitioners hit when their model is just large enough to memorize the training set but not yet large enough to find the smooth generalizing solution.

The mechanism is bias-variance breakdown near the interpolation threshold. Just below it, the model has to use all its parameters to fit the training set, so any small change in the data forces a large change in the parameters — variance is enormous. Just above it, there are many parameter configurations that fit the training set, and gradient descent finds one that lies in a low-variance basin of the loss landscape. The classical U-curve does not predict the second descent because it assumes a unique minimizer for each dataset, which is false in the overparameterized regime.

This is why the modern recipe is: pick a model much larger than the classical theory suggests, train it with strong regularization (weight decay, dropout, data augmentation), and stop when the validation loss stops improving. The model is in the second descent regime; the regularization keeps the iterate out of the high-variance basin near the interpolation threshold.

Dropout: a Bayesian approximation

Dropout trains an ensemble of networks in parallel: at each step, every neuron is independently zeroed out with probability pp. At test time, all neurons are used, but the weights are scaled by 1/(1−p)1/(1-p) to keep the expected activation constant. The result is that the prediction is the geometric mean of an exponential number of "thinned" networks, which acts as a regularizer.

The Bayesian interpretation is that dropout is approximately variational inference in a Bernoulli belief network: each weight has a posterior distribution over "active" and "inactive", and dropout samples from that distribution at training time. The expected prediction under the dropout distribution is what the test-time scaling approximates.

Dropout is one form of regularization among many. For convolutional layers, it has largely been replaced by batch normalization (which has its own regularization effect through the noise in the minibatch statistics). For transformers, weight decay (AdamW) and stochastic depth (dropping entire layers) are the standard remedies. For tabular data, L2 and early stopping remain the defaults.

def dropout_forward(x, p, training=True):
    """Inverted dropout: scale by 1/(1-p) at training time, identity at test time."""
    if not training:
        return x
    mask = (np.random.rand(*x.shape) >= p) / (1 - p)
    return x * mask

The interaction between dropout and the optimizer is subtle. Dropout noise is not independent across steps — the same neuron can be active or inactive in a way that depends on the data and the previous activations — so it does not act like i.i.d. Gaussian noise on the gradient. The effective regularization strength is therefore not a simple function of pp; you tune pp on a validation set, not by theory.

Elastic net: what the mixture actually buys you

Suppose two features are near-duplicates — say income and income_log — so their design-matrix columns are almost collinear. Lasso's geometry then does something unpleasant: the L1 ball touches one axis, the objective is nearly flat along the direction that trades the two features off against each other, and coordinate descent picks an essentially arbitrary one of them, dropping the other. The model is sparse, but it is unstable sparse — re-run on a resampled dataset and a different feature survives.

The fix is not to abandon sparsity but to add a little L2 back:

w^EN  =  arg⁡min⁡w[12n∥y−Xw∥22  +  α∥w∥1  +  12(1−α)∥w∥22]\hat{w}_{\text{EN}} \;=\; \arg\min_w \Bigl[ \tfrac{1}{2n} \lVert y - Xw \rVert_2^2 \;+\; \alpha \lVert w \rVert_1 \;+\; \tfrac{1}{2}(1-\alpha) \lVert w \rVert_2^2 \Bigr]

The ∥w∥22\lVert w \rVert_2^2 term makes the objective strictly convex and smooth, so the "nearly flat ridge" between correlated features is removed and the selected set becomes stable. α\alpha then interpolates: α=1\alpha = 1 is pure lasso, α=0\alpha = 0 is pure ridge. What you have added is not a second regularizer for its own sake — it is a nonsmoothness remedy, and its price is that a genuinely useless feature can no longer be pushed to exactly zero.

The path is worth plotting. Fit the model over a grid of α\alpha values and record which features enter; features that are genuinely causal enter at every α\alpha and form a horizontal band, while features that enter only at α\alpha near 1 are collinear artifacts. In sklearn that is one line — ElasticNet(l1_ratio=...).path(x, y) — and it is the closest thing to a stability-selection bootstrap that runs in seconds.

from sklearn.linear_model import ElasticNet
import numpy as np

alphas = np.logspace(-4, 0, 40)              # sweep the overall strength
enet = ElasticNet(l1_ratio=0.5, max_iter=5000)
coefs = enet.path(X_train, y_train, coef_init=np.zeros(X_train.shape[1]),
                  alpha=alphas / alphas.max())[1]
# Columns that are non-zero for the whole sweep are the stable ones.

# Elastic net has no closed form, but it does have a coordinate-descent update that is
# the soft-threshold operator with a ridge correction:
def enet_coordinate_update(z, G_jj, alpha, l1_ratio, lam):
    """Minimise  0.5*G_jj*w_j^2 + l1_ratio*lam*|w_j| + 0.5*(1-l1_ratio)*lam*w_j^2
    holding all other coordinates fixed. The penalty is quadratic-plus-L1, so the
    solution is a soft-threshold scaled by the curvature the penalty itself adds.
    """
    penalty = l1_ratio * lam
    ridge = (1 - l1_ratio) * lam
    # Effective curvature is the data curvature PLUS the ridge curvature.
    return np.sign(z) * max(abs(z) - penalty, 0.0) / (G_jj + ridge)

Why L2 and AdamW are not the same thing

This is the single most common confusion carried over from ML-101, so it is worth being precise. With plain SGD, L2 regularization is implemented by adding the penalty to the gradient:

gregularized=∇L+2λwg_{\text{regularized}} = \nabla \mathcal{L} + 2\lambda w

and the weight therefore decays at a rate proportional to its own magnitude — a multiplicative shrinkage. But add momentum or Adam and that equivalence breaks. Adam normalises the gradient by its second moment:

mt=β1mt−1+(1−β1)gt,vt=β2vt−1+(1−β2)gt2m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t, \qquad v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2

wt=wt−1−η m^tv^t+ϵw_t = w_{t-1} - \eta \,\frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon}

Now put the L2 term into gtg_t. For a weight whose only gradient signal is the penalty, gt=2λwg_t = 2\lambda w, so mt/vt≈sign⁡(w)m_t/\sqrt{v_t} \approx \operatorname{sign}(w) — the magnitude of ww cancels out completely. The update becomes a constant-size step toward zero, independent of how large the weight is. Coupling the penalty through Adam has therefore silently converted proportional shrinkage into sign-like shrinkage, which decays small weights far too slowly and changes the effective regularizer as training proceeds.

AdamW fixes this by moving the decay out of the gradient and into the parameter update:

wt=wt−1−η m^tv^t+ϵ  −  ηλwt−1w_t = w_{t-1} - \eta \,\frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} \;-\; \eta\lambda w_{t-1}

The decay term now sits beside the Adam step rather than inside it, so it is not divided by v^t\sqrt{\hat{v}_t} and retains its multiplicative character. The practical symptom of getting this wrong: on a transformer fine-tune, L2-in-Adam under-regularises small weights, the model fails to reach the sparsity you expected, and a sweep over weight_decay produces a flat, confusing response.

Data augmentation as a prior over invariances

Augmentation is usually taught as a trick to manufacture more data. At 201 depth it is better read as an explicit statement of which transformations of the input you believe leave the label unchanged — a prior over invariances, in the same sense that L2 is a prior over weight magnitude.

Under the invariance hypothesis, if gg is a transformation with y(g(x))=y(x)y(g(x)) = y(x), then the correctly inflated training distribution is the one that marginalises over the group GG generated by gg. The model that minimises risk over that distribution is the one whose predictions are constant along group orbits. So augmentation is a way of encoding a symmetry into the hypothesis class rather than hoping SGD discovers it.

AugmentationInvariance assertedFails when
Horizontal flipCamera orientation is irrelevantThe class encodes orientation (digits, traffic side)
Small translationObject position is nuisanceAbsolute position carries the label (geography, layout)
Colour jitterLighting is nuisanceColour is the signal (disease grading, spectrograms)
CutMix / MixUpConvexity of the data manifoldLabels are not linear in the inputs
Specular / time masksOcclusion is transientThe occluded region is systematically informative

That last column is the point. Every augmentation is a bet, and a bet is falsifiable. MixUp is not a free regularizer; its implicit assumption is that the data manifold is locally convex and that labels blend linearly. On a task where the classes are defined by a conjunction of features — "has both property A and property B" — a half-and-half image has no meaningful label and MixUp injects label noise. This is why augmentation ablations on a specific dataset are worth more than any prior belief about which one is better.

Label smoothing and what it really assumes

Label smoothing replaces the one-hot target yy with a mixture of the one-hot vector and the uniform distribution:

y~=(1−ε) y  +  εK1\tilde{y} = (1-\varepsilon)\, y \;+\; \frac{\varepsilon}{K} \mathbf{1}

Minimising cross-entropy against y~\tilde{y} pushes the model's predicted distribution toward the smoothed target rather than the exact one. The 101-level story is "it stops the model being overconfident". The precise story is what it does to the logit gap.

Let z1z_1 be the logit of the true class and max⁡k≠1zk\max_{k \ne 1} z_k the largest competitor. With no smoothing, the loss is unbounded below and the optimum is z1−max⁡kzk→∞z_1 - \max_k z_k \to \infty: the network is rewarded for driving the correct logit to ever-higher values even after the prediction is certainly right. Smoothing bounds it. Set ∂L/∂zk=0\partial \mathcal{L}/\partial z_k = 0 at the smoothed optimum and you get

z1∗−max⁡k≠1zk∗=log⁡ ⁣((1−ε)(K−1)ε+1)z_1^\ast - \max_{k \ne 1} z_k^\ast = \log\!\Bigl(\frac{(1-\varepsilon)(K-1)}{\varepsilon} + 1\Bigr)

which is finite and, for ε=0.1\varepsilon = 0.1 and K=1000K = 1000, is about log⁡(8992)≈9.1\log(8992) \approx 9.1. So label smoothing is a bounded-confidence prior: it caps how much probability mass the model may place on its most likely class, and the cap is a deterministic function of ε\varepsilon and KK. Two consequences that matter in practice. First, the mechanism is temperature-like, so it interacts with a temperature-scaled softmax at inference — you should not apply both blindly. Second, it is a prior on the output distribution, not a data augmentation, so it is entirely orthogonal to the input-side priors above; combining one does not substitute for the other.

Choosing a regularizer: a decision procedure

The order below is not arbitrary — it goes from assumptions you already believe, to knobs that cost the most to search.

  1. Start with none. If training error and test error are both high, the problem is underfitting and no regularizer will help. Adding one here makes it worse.
  2. Add early stopping. It is free, needs no new hyperparameter beyond the patience you already tuned, and it is the closest thing to a default.
  3. Add weight decay in its decoupled form (AdamW, or SGD with momentum where the equivalence does hold). This is the cheap, well-understood default for any model trained with an adaptive optimizer.
  4. Add data augmentation if you have a real invariance story. It is the only entry here that can improve the bias term rather than trading it away.
  5. Reach for L1/elastic net only when you specifically need variable selection on a tabular problem, and verify stability by re-running on resampled data.
  6. Use dropout for dense architectures where weight decay underperforms, and skip it entirely for transformers, where the standard stack is weight decay plus stochastic depth plus label smoothing.
def diagnose_bias_variance(X_tr, y_tr, X_va, y_va, factory, n_seeds=10):
    """
    The only honest way to choose a regularizer: measure the two terms separately by
    refitting on independent draws. A single train/validate split conflates them; this
    averages the estimator's variance across resamples and exposes its bias as the gap
    between the mean prediction and the validation truth.
    """
    preds = np.array([factory().fit(X_tr, y_tr).predict(X_va) for _ in range(n_seeds)])
    mean_pred = preds.mean(axis=0)                  # E_D[f(x; w_hat)] over resamples
    bias_sq   = ((mean_pred - y_va) ** 2).mean()
    variance  = preds.var(axis=0, ddof=1).mean()    # Var_D[f(x; w_hat)]
    noise     = ((y_va - mean_pred) ** 2).mean() - variance   # crude sigma^2 estimate
    return {"bias^2": bias_sq, "variance": variance,
            "irreducible": max(0.0, noise), "total": bias_sq + variance + max(0.0, noise)}

The resampling matters. Choosing a regularizer on a single validation split and then reporting that same split's error is the same mistake as tuning on the test set, and it is the reason "my validation curve looks U-shaped" is weak evidence for anything. The decomposition above needs independent draws to be meaningful at all.

Key Takeaways

  • L2 regularization is the MAP estimate under a Gaussian prior on the weights; the regularization coefficient is the prior-precision ratio σ2/τ2\sigma^2 / \tau^2, not a free magic constant. L1 is the MAP under a Laplace prior and produces sparse solutions because the L1 ball has corners that touch coordinate axes.
  • The bias-variance decomposition explains the classical U-curve: regularizers trade bias for variance, and the optimum is where the sum is maximised. This decomposition assumes an underparameterized model with a unique minimizer.
  • Double descent is the failure of that assumption: in the overparameterized regime, the test-error curve is non-monotone, with a peak at the interpolation threshold. Modern deep learning lives in the second-descent regime; strong regularization keeps the iterate out of the high-variance basin.
  • Early stopping is implicit per-eigendirection L2 shrinkage: directions of high curvature are shrunk fast, directions of low curvature are shrunk slowly. This is why early stopping is sometimes a substitute for explicit regularization.
  • Dropout, weight decay, batch normalization, and stochastic depth are all regularizers with different assumptions. The choice depends on the prior belief about the model: dropout for overconfident deep nets, weight decay for transformers, batch normalization for convolutional architectures.

Check your understanding

6 questions · 80% to complete the lesson

1 / 6

5 correct to pass

L2 regularization is the MAP estimate under a Gaussian prior on the weights. The regularization coefficient lambda is therefore which quantity?

0 of 6 answered

Pick a lesson to start the audio.