03

Designing Loss Functions

Cantonese podcast title: 損失函數的設計

Learning Objectives

  1. State the negative log-likelihood of the data under a noise model, and show that
  2. Derive squared-error, cross-entropy, hinge, Huber, and focal losses from their
  3. Identify the *outlier robustness* property of the Huber loss and the
  4. Diagnose a mis-specified loss by checking whether the residuals are symmetric, or
  5. Translate a noise model into the corresponding PyTorch-style loss function call,
Designing Loss Functions — visual guide
Loss functions and their residual profiles A plot of five loss functions versus residual r = y - y_hat. Squared error grows quadratically; absolute error grows linearly; Huber is quadratic for small r and linear for large r; hinge is zero for large-margin predictions and linear for margin violations; cross-entropy (shown for log-odds) grows super-linearly for wrong-sign predictions. ℓ(r) r 0 -3 -1.5 +1.5 +3 10 5 squared absolute Huber δ=1.5 hinge (V-shape) Loss functions vs residual Each curve is a different assumption about the noise model Noise model Gaussian (constant variance) Laplace (heavy tail)

Assumes you know from ML-101

You have completed ML-101 lesson 4 — Logistic Regression, where you saw the cross-entropy loss and were told it is the right loss for binary classification; ML-101 lesson 9 — Model Evaluation, where hinge loss appeared without derivation in the SVM section; and ML-101 lesson 14 — Neural Networks & Backprop, where mean-squared error was used as the headline regression loss. This lesson does not re-derive the logistic function, does not re-state the chain rule, and does not re-teach what a probability is. It assumes you already know the names of the losses; it derives them from the noise model they imply, and asks: given a prediction problem, which noise model is the right one, and what is the cost of choosing the wrong one.

Learning Objectives

  1. State the negative log-likelihood of the data under a noise model, and show that the resulting loss is determined by the noise model — not by "how much we want to punish errors".
  2. Derive squared-error, cross-entropy, hinge, Huber, and focal losses from their respective noise models, and name the regime in which each is Bayes-consistent.
  3. Identify the outlier robustness property of the Huber loss and the foreground-vs-background reweighting property of focal loss, and connect each to the noise model that produces it.
  4. Diagnose a mis-specified loss by checking whether the residuals are symmetric, or by checking whether the calibrated probability of a prediction matches its observed frequency in a held-out set.
  5. Translate a noise model into the corresponding PyTorch-style loss function call, and predict which term in the loss corresponds to which assumption.

The negative log-likelihood derivation

The bridge from "what noise model did the data come from" to "what loss should I minimise" is the negative log-likelihood. Given a parameterised model pθ(y∣x)p_\theta(y \mid x) and a dataset D={(xi,yi)}i=1n\mathcal{D} = \{(x_i, y_i)\}_{i=1}^n assumed i.i.d. from the true joint p⋆(x,y)p^\star(x, y), the log-likelihood is

log⁡L(θ)  =  ∑i=1nlog⁡pθ(yi∣xi).\log \mathcal{L}(\theta) \;=\; \sum_{i=1}^{n} \log p_\theta(y_i \mid x_i).

The maximum-likelihood estimator is

θ^MLE  =  arg⁡max⁡θlog⁡L(θ)  =  arg⁡min⁡θ∑i=1n−log⁡pθ(yi∣xi),\hat{\theta}_{\mathrm{MLE}} \;=\; \arg\max_{\theta} \log \mathcal{L}(\theta) \;=\; \arg\min_{\theta} \sum_{i=1}^{n} -\log p_\theta(y_i \mid x_i),

so the loss function is, by definition,

ℓ(y,p^)  =  −log⁡pθ(y∣x),\ell(y, \hat{p}) \;=\; -\log p_\theta(y \mid x),

where p^\hat{p} is whatever statistic of θ\theta and xx the noise model requires (typically the conditional mean, the conditional probability of a class, or the parameter of a count distribution). The derivation has no choice: a noise model is a loss, and a loss is a noise model in disguise.

The deeper statement is that the maximum-likelihood estimator is Bayes-optimal under the assumed noise model. If the true data-generating distribution is p⋆(x,y)≠pθ(y∣x)pθ(x)p^\star(x, y) \neq p_\theta(y \mid x) p_\theta(x) for any θ\theta, then no maximum-likelihood estimator is Bayes-optimal, and minimising the resulting loss will not produce a calibrated predictor. The cost of mis-specification is not a scalar; it is the asymptotic gap between the chosen predictor and the true conditional p⋆(y∣x)p^\star(y \mid x), which the Bernstein–von Mises theorem bounds but does not eliminate.

The negative log-likelihood is also the only loss that is proper in the strict sense: averaged over the data, it is uniquely minimised at the true distribution, not at any surrogate. Squared error and cross-entropy are both proper because they are negative log-likelihoods of the Gaussian and categorical families. Hinge loss and Huber loss are not negative log-likelihoods of any exponential family; they are surrogate losses that are Bayes-consistent — meaning their minimiser over the function class coincides with the Bayes classifier — but not proper. The distinction matters when the loss is used for calibration, not just for classification accuracy.

import numpy as np

def neg_log_likelihood(log_p: np.ndarray) -> float:
    """Generic NLL: -sum log p_theta(y_i | x_i). Operates on log-probabilities.

    Why log-probabilities rather than probabilities: log(p) is well-behaved
    underflows; for p = 1e-300 the log is finite, but p itself rounds to zero
    in float64, which makes -log(p) silently infinite in any downstream check.
    """
    return float(-log_p.sum())

# Reading: log_p[i] = log p_theta(y_i | x_i).  NLL is just the negation.
log_p = np.array([-1.2, -0.4, -2.1, -0.7])  # toy per-example log-probs
print("NLL:", round(neg_log_likelihood(log_p), 4))

Squared error from Gaussian noise

Suppose y∣x∼N(fθ(x),σ2)y \mid x \sim \mathcal{N}(f_\theta(x), \sigma^2) for some fixed variance σ2\sigma^2. Then

log⁡pθ(y∣x)  =  −12σ2(y−fθ(x))2+const.\log p_\theta(y \mid x) \;=\; -\frac{1}{2\sigma^2}\bigl(y - f_\theta(x)\bigr)^2 + \text{const}.

The constant absorbs the log normaliser −12log⁡(2πσ2)-\tfrac{1}{2}\log(2\pi\sigma^2), which does not depend on θ\theta. The negative log-likelihood is, up to a positive multiplicative constant,

ℓ(y,y^)  =  (y−fθ(x))2.\ell(y, \hat{y}) \;=\; \bigl(y - f_\theta(x)\bigr)^2.

Squared error is the negative log-likelihood under Gaussian noise with fixed variance. The variance σ2\sigma^2 scales the loss but does not change the minimiser — under constant variance, dividing the loss by 2σ22\sigma^2 is a trivial rescaling.

The deeper point is the variance assumption. A loss derived under σ2=σ2(x)\sigma^2 = \sigma^2(x) — heteroscedastic Gaussian noise — gives

ℓ(y,y^)  =  (y−fθ(x))2σ2(x),\ell(y, \hat{y}) \;=\; \frac{\bigl(y - f_\theta(x)\bigr)^2}{\sigma^2(x)},

which is a weighted squared error with weights 1/σ2(x)1/\sigma^2(x). The 101-level model "use MSE everywhere" is the special case where the noise variance is assumed constant; the 201-level reading is that the assumption of constant variance is stronger than the assumption of Gaussian noise, and it fails on every real regression dataset where the residual scale depends on xx.

The 201-level test of whether the noise is Gaussian is the distribution of the residuals. A Q–Q plot of residuals against the standard normal should fall on a line; heavy tails produce Q–Q plots that bend at the extremes, and a maximum-likelihood estimator under Cauchy or Laplace noise (the next sections) recovers better predictions. The diagnosis is data-driven, not theoretical.

Cross-entropy from categorical noise

For a KK-class classification problem, suppose y∣x∼Categorical(σ(fθ(x)))y \mid x \sim \mathrm{Categorical}(\sigma(f_\theta(x))), where σ\sigma is the softmax that maps Rm\mathbb{R}^m to the KK-simplex. Then

log⁡pθ(y∣x)  =  log⁡σy(fθ(x))  =  fθ(x)y−log⁡∑k=1Kexp⁡fθ(x)k.\log p_\theta(y \mid x) \;=\; \log \sigma_y(f_\theta(x)) \;=\; f_\theta(x)_y - \log\sum_{k=1}^{K} \exp f_\theta(x)_k.

The negative log-likelihood is the cross-entropy of the one-hot label yy with the predicted softmax distribution:

ℓ(y,p^)  =  −∑k=1Kyklog⁡p^k.\ell(y, \hat{p}) \;=\; - \sum_{k=1}^{K} y_k \log \hat{p}_k.

For binary classification this reduces to −(ylog⁡p^+(1−y)log⁡(1−p^))-\bigl(y \log \hat{p} + (1 - y)\log(1 - \hat{p})\bigr), the logistic-regression loss from ML-101 lesson 4.

The deeper point is why the softmax link is the right one. The softmax σ(z)k=exp⁡zk/∑jexp⁡zj\sigma(z)_k = \exp z_k / \sum_j \exp z_j is the normaliser of the exponential family whose natural parameter is zz and whose sufficient statistic is the one-hot vector. The log-partition function A(z)=log⁡∑kexp⁡zkA(z) = \log \sum_k \exp z_k is computable in closed form, which is why the negative log-likelihood is tractable. A different link function — sigmoid for binary, probit, or any non-canonical link — produces a likelihood that does not have a closed-form normaliser in general. The softmax is special because the exponential family with the categorical sufficient statistic has a closed-form A(z)A(z), not because "softmax is a smooth approximation to argmax".

import numpy as np

def softmax(z: np.ndarray) -> np.ndarray:
    # Numerically stable softmax: subtract the max for log-sum-exp.
    z = z - z.max(axis=-1, keepdims=True)
    e = np.exp(z)
    return e / e.sum(axis=-1, keepdims=True)

def cross_entropy(y_onehot: np.ndarray, p: np.ndarray) -> float:
    # Per-example NLL under categorical noise. y_onehot: (n, K), p: (n, K).
    p = np.clip(p, 1e-12, 1.0)              # log(0) is -inf; clip keeps the loss finite.
    return float(-(y_onehot * np.log(p)).sum(axis=-1).mean())

# Toy: 100 examples, 3 classes, model outputs pre-softmax logits.
rng = np.random.default_rng(0)
logits = rng.standard_normal((100, 3))
y      = np.eye(3)[rng.integers(0, 3, 100)]   # one-hot labels
p      = softmax(logits)
print("mean cross-entropy:", round(cross_entropy(y, p), 4))

Hinge loss and the label-noise interpretation

The hinge loss is the surrogate used by support vector machines:

ℓ(y,y^)  =  max⁡(0, 1−y⋅y^),\ell(y, \hat{y}) \;=\; \max(0,\, 1 - y \cdot \hat{y}),

where y∈{−1,+1}y \in \{-1, +1\} and y^=fθ(x)\hat{y} = f_\theta(x). The hinge is not the negative log-likelihood of any exponential family, but it is Bayes-consistent under the following interpretation: the label is the true class with probability 1−η1 - \eta, and is flipped uniformly to the opposite class with probability η\eta, where η<0.5\eta < 0.5. Under that noise model, the Bayes-optimal classifier is the sign of fθ(x)f_\theta(x), and the hinge loss is the convex surrogate that upper-bounds the 0/1 loss and admits a unique finite minimiser.

The deeper point is why the hinge is convex. Squared error and cross-entropy are convex in fθf_\theta because they are negative log-likelihoods of exponential-family models with convex log-partition functions. Hinge is convex because it is the pointwise maximum of affine functions, which is convex by construction. Convexity is the property that gradient descent converges; Bayes-consistency is the property that the minimiser is correct. A loss can be convex without being Bayes-consistent (e.g., the calibration is wrong) and Bayes-consistent without being convex (e.g., the 0/1 loss itself).

The 101-level mistake is to treat the hinge loss as a one-off SVM curiosity. The 201-level reading is that the hinge is the convex surrogate for the symmetric label-noise model, and the symmetric assumption — flipping to the opposite class with equal probability — is the place where the result breaks. Asymmetric label noise requires a different surrogate, and the cross-entropy is the canonical surrogate for no label noise, since the categorical model already absorbs the noise into the conditional class probability.

Huber loss and outlier-robust estimation

The Huber loss interpolates between squared error (for small residuals) and absolute error (for large residuals):

ℓ(y,y^)  =  {12(y−y^)2if ∣y−y^∣≤δ,δ∣y−y^∣−12δ2otherwise.\ell(y, \hat{y}) \;=\; \begin{cases} \tfrac{1}{2}(y - \hat{y})^2 & \text{if } \lvert y - \hat{y} \rvert \le \delta, \\ \delta \lvert y - \hat{y} \rvert - \tfrac{1}{2}\delta^2 & \text{otherwise}. \end{cases}

The Huber loss is the negative log-likelihood under a contaminated Gaussian noise model: with probability 1−ε1 - \varepsilon, y∣x∼N(fθ(x),σ2)y \mid x \sim \mathcal{N}(f_\theta(x), \sigma^2), and with probability ε\varepsilon, y∣x∼N(fθ(x),c2σ2)y \mid x \sim \mathcal{N}(f_\theta(x), c^2\sigma^2) for some large cc. The mixing of a "narrow" and a "wide" Gaussian produces a distribution with heavier tails than the Gaussian, and the maximum likelihood estimator under the mixture is the Huber minimiser in the limit.

The deeper point is that the Huber loss is robust in the precise sense: the influence function of the M-estimator is bounded, so a single contaminated example cannot move the fit by more than a constant. Squared error has an unbounded influence function, so a single outlier with ∣y−y^∣=100|y - \hat{y}| = 100 has an influence 100 times larger than a residual of size 1. The Huber loss caps that influence at δ\delta, at the cost of being non-differentiable at the kink.

The 101-level reading "Huber is robust to outliers" is true but vague. The 201-level reading is that robustness is a property of the influence function, and the Huber loss is the maximum-likelihood estimator under a contaminated Gaussian noise model with contamination fraction ε\varepsilon. Choosing δ\delta is related to ε\varepsilon and the noise variance σ2\sigma^2 by an explicit equation, and the practitioner who treats δ\delta as a free knob is using the loss without using the noise model.

import numpy as np

def huber(y: np.ndarray, yhat: np.ndarray, delta: float = 1.0) -> float:
    r = np.abs(y - yhat)
    quad = 0.5 * r ** 2
    lin  = delta * r - 0.5 * delta ** 2
    return float(np.where(r <= delta, quad, lin).mean())

# Toy: one outlier at y - yhat = 100.
rng = np.random.default_rng(0)
yhat = np.zeros(100)
y    = rng.standard_normal(100)
y[0] = 100.0
print("Huber  :", round(huber(y, yhat, delta=1.0), 4))
print("MSE    :", round(((y - yhat) ** 2).mean() / 2, 4))   # squared error, scaled

Focal loss and foreground reweighting

The focal loss, introduced for one-stage object detection, is

ℓ(y,p^)  =  −(1−p^t)γlog⁡p^t,\ell(y, \hat{p}) \;=\; -(1 - \hat{p}_t)^\gamma \log \hat{p}_t,

where p^t=p^\hat{p}_t = \hat{p} if y=1y = 1 and p^t=1−p^\hat{p}_t = 1 - \hat{p} otherwise, and γ≥0\gamma \ge 0 is a focusing parameter. At γ=0\gamma = 0, focal loss reduces to cross-entropy. For γ>0\gamma > 0, examples that the model already classifies correctly with high confidence have p^t\hat{p}_t close to 1, so the factor (1−p^t)γ(1 - \hat{p}_t)^\gamma is small and their contribution to the gradient shrinks. Examples that the model misclassifies have p^t\hat{p}_t small, the factor is close to 1, and their contribution to the gradient is preserved.

The deeper point is that focal loss is not a noise model; it is a reweighting of the cross-entropy loss that compensates for an asymmetric prior on the training distribution. The noise model underlying focal loss is the same as for cross-entropy — categorical — but the empirical class frequencies in the training data are heavily skewed (e.g., 1000 background examples per foreground object in detection). The maximum-likelihood estimator under categorical noise minimises the weighted cross-entropy where the weights are the inverse of the class frequencies; focal loss approximates that reweighting by a smooth, confidence-dependent factor.

The 101-level reading "focal loss is for imbalanced data" is true. The 201-level reading is that focal loss is a smooth surrogate for inverse-frequency reweighting, and the optimal γ\gamma depends on the imbalance ratio and the classifier's confidence distribution. The link between γ\gamma and the imbalance ratio is approximate; in practice γ\gamma is tuned as a hyperparameter.

How to choose a loss in practice

The decision tree for choosing a loss is short, but each branch requires the practitioner to commit to an assumption about the data.

LossNoise modelWhen to choose
Squared errorConstant-variance GaussianRegression with symmetric, light-tailed residuals
Cross-entropyCategoricalMulti-class or binary classification with calibrated probabilities as output
HingeSymmetric label flippingHard-margin classification, asymmetric cost not required
HuberContaminated GaussianRegression with occasional outliers that should not dominate the fit
FocalCategorical + inverse-frequency reweightingClassification with severe class imbalance and many "easy" negatives
Quantile (pinball)Asymmetric LaplaceRegression with asymmetric cost; e.g., demand at risk of stock-out

The 201-level mistake is to choose a loss for computational convenience rather than for the noise model. Cross-entropy is the easy loss to optimise because the softmax link makes the gradient well-scaled; it is not automatically the right loss for every classification problem. Squared error is the easy regression loss to backprop; it is not automatically the right loss for every regression problem.

The 201-level test is calibration. If the loss is the correct negative log-likelihood for the data, the predicted probabilities from a held-out set should match the observed frequencies — a reliability diagram should fall on the diagonal. A loss that produces mis-calibrated probabilities is either mis-specified or under-trained, and the diagnostic distinguishes the two. Cross-entropy minimisation can under-train and leave probabilities mis-calibrated; squared error on a classification problem produces mis-calibrated probabilities even when fully trained, because the Gaussian noise model is wrong.

The choice of loss is therefore a modelling decision, not an optimisation decision. ML-101 lesson 4 said "use cross-entropy for classification"; ML-201 lesson 3 says "use the negative log-likelihood of the noise model you believe generated the data, and verify your belief with a calibration plot on a held-out set before deploying".

Key Takeaways

  • The negative log-likelihood is the bridge from a noise model to a loss. A noise model is a loss in disguise; a loss is a noise model in disguise.
  • Squared error is the loss for constant-variance Gaussian noise; cross-entropy is the loss for categorical noise; Huber is the loss for contaminated Gaussian noise; hinge is the convex surrogate for symmetric label-flipping noise; focal is cross-entropy reweighted by an inverse-frequency factor.
  • The 201-level test of a correct loss is calibration on a held-out set, not the value of the training loss. A small training loss with poor calibration is a mis-specified loss, not a well-trained model.
  • The choice of loss is a modelling decision, not an optimisation decision. The gradient-descent routine does not care which loss you pick; the generative behaviour of the deployed model does.

Check your understanding

7 questions · 80% to complete the lesson

1 / 7

6 correct to pass

Squared error ℓ(y,y^)=(y−y^)2\ell(y, \hat{y}) = (y - \hat{y})^2 is the negative log-likelihood of the data under which noise model?

0 of 7 answered

Pick a lesson to start the audio.