09

Imbalanced Data

Cantonese podcast title: 不平衡數據

Learning Objectives

  1. Derive the Bayes-optimal decision threshold under unequal misclassification costs and unequal class priors, and explain why the textbook 0.5 is a special case rather than the general one.
  2. Compute precision, recall, F1, PR-AUC and ROC-AUC from the same scored classifier and explain why PR-AUC is the more honest summary on a heavily imbalanced test set.
  3. Implement cost-sensitive learning by class weighting, by threshold tuning, and by resampling, and state the assumption whose violation makes each technique silently worse than doing nothing.
  4. Implement SMOTE and its two failure modes (over-sampling noise and bleeding across the class boundary), and explain why borderline-SMOTE and Tomek links exist.
  5. Diagnose a deployed classifier whose overall accuracy is high but whose positive-class recall is collapsing, using only its score distribution and confusion matrix.
Imbalanced Data — visual guide
ROC vs PR curves on the same scored classifier Same scores, two summaries: ROC vs PR on 1:99 priors AUC can look great on ROC while precision collapses on PR ROC: FPR on x, TPR on y false positive rate true positive rate 1.0 1.0 AUC 0.96 - looks great random diagonal 0.0 0.0 PR: recall on x, precision on y PR-AUC ~0.30 - precision dies recall precision 1.0 1.0 prior 0.01 0.0 ROC counts TN; PR does not. On 1:99 priors TN dominates the x-axis and ROC cannot see the precision collapse. Bayes-optimal threshold tau* = C_01 / (C_01 + C_10); 0.5 only when C_01 = C_10.

Assumes you know from ML-101

This lesson builds on ML-101 Lesson 9 (Model Evaluation) and ML-101 Lesson 4 (Logistic Regression). The reader is presumed to already know how a confusion matrix is built from a thresholded classifier, what precision and recall mean individually, that accuracy is the diagonal sum over the total, and that logistic regression produces a probability p^=σ(z)\hat{p} = \sigma(z) that you compare against 0.5 by default. The reader is also expected to know what a "minority class" is at the column-counting level — that the label histogram is not 50/50 and that a degenerate "always-predict-the-majority" rule has accuracy equal to the majority proportion.

What ML-101 did not do is ask why 0.5 is the default threshold, whether AUC is the right metric to report, or how to alter training itself when the class prior is hostile. This lesson derives all three from the same place — Bayes risk minimisation — and shows which assumption fails when the wrong tool is picked.

Learning Objectives

  1. Derive the Bayes-optimal decision threshold under unequal misclassification costs and unequal class priors, and explain why the textbook 0.5 is a special case rather than the general one.
  2. Compute precision, recall, F1, PR-AUC and ROC-AUC from the same scored classifier and explain why PR-AUC is the more honest summary on a heavily imbalanced test set.
  3. Implement cost-sensitive learning by class weighting, by threshold tuning, and by resampling, and state the assumption whose violation makes each technique silently worse than doing nothing.
  4. Implement SMOTE and its two failure modes (over-sampling noise and bleeding across the class boundary), and explain why borderline-SMOTE and Tomek links exist.
  5. Diagnose a deployed classifier whose overall accuracy is high but whose positive-class recall is collapsing, using only its score distribution and confusion matrix.

The Bayes-optimal decision rule

Take a scored classifier that outputs p^(y=1∣x)\hat{p}(y=1 \mid x). For a single test point, two errors are possible:

  • predict 0 when the truth is 1 — a false negative,
  • predict 1 when the truth is 0 — a false positive.

Let C10C_{10} be the cost of predicting 0 when y=1y=1 and C01C_{01} the cost of predicting 1 when y=0y=0. The expected cost of predicting y^=1\hat{y}=1 on this point is

E[C∣y^=1]  =  C10 Pr⁡(y=1∣x).\mathbb{E}[C \mid \hat{y}=1] \;=\; C_{10}\,\Pr(y=1 \mid x).

The expected cost of predicting y^=0\hat{y}=0 is

E[C∣y^=0]  =  C01 Pr⁡(y=0∣x).\mathbb{E}[C \mid \hat{y}=0] \;=\; C_{01}\,\Pr(y=0 \mid x).

The classifier should predict 1 iff the cost of doing so is smaller:

C10 Pr⁡(y=1∣x)  <  C01 Pr⁡(y=0∣x).C_{10}\,\Pr(y=1 \mid x) \;<\; C_{01}\,\Pr(y=0 \mid x).

Substitute Pr⁡(y=0∣x)=1−Pr⁡(y=1∣x)\Pr(y=0 \mid x) = 1 - \Pr(y=1 \mid x) and rearrange:

p^(y=1∣x)  >  τ⋆  ≡  C01C01+C10.\hat{p}(y=1 \mid x) \;>\; \tau^{\star} \;\equiv\; \frac{C_{01}}{C_{01} + C_{10}}.

This is the Bayes-optimal threshold. It does not depend on the model's class balance — it depends only on the cost ratio. The textbook default τ=0.5\tau=0.5 is recovered only when C01=C10C_{01} = C_{10}, i.e. when the two errors hurt equally. The moment one error hurts more than the other, 0.5 is wrong on every prediction the classifier makes.

What does change with the prior is the unconditioned error rate: a 1:99 prior means most points are negative regardless of xx, and the frequency of errors depends on Pr⁡(y=1)\Pr(y=1). But the rule at every xx depends only on the costs. Confusion between these two facts is the source of most imbalanced-data advice on the internet.

import numpy as np

def bayes_threshold(cost_fn: float, cost_fp: float) -> float:
    """
    Bayes-optimal threshold for class-1 under unequal misclassification costs.

    Cost of a false negative (miss a real positive) is `cost_fn`; cost of a
    false positive (false alarm) is `cost_fp`. Returns the probability above
    which the classifier should output 1.
    """
    return cost_fp / (cost_fp + cost_fn)


# Example: catching fraud. Missing one fraud costs $1000 in chargebacks;
# blocking a legitimate customer costs $50 in support time and churn.
tau = bayes_threshold(cost_fn=1000.0, cost_fp=50.0)
print(f"tau* = {tau:.4f}")  # tau* = 0.0476 — predict fraud above 4.76% score

The output is τ⋆≈0.0476\tau^\star \approx 0.0476. Anything the model scores above 4.76% is more expensive to miss than to alarm on, so the operating threshold is two orders of magnitude below the textbook default. This is why fraud models that "predict fraud at 0.5" never see a single alert.

Why accuracy is the wrong metric on skewed priors

Accuracy on a 1:99 prior with a model that always outputs 0 is

accdegenerate  =  0.99.\mathrm{acc}_{\text{degenerate}} \;=\; 0.99.

That number is high, accurate, and useless: it tells you the classifier has learned nothing about the positive class. Accuracy is a weighted average of per-class correct rates, with weights equal to the class priors:

acc  =  π1 TPR  +  π0 TNR,\mathrm{acc} \;=\; \pi_{1}\,\mathrm{TPR} \;+\; \pi_{0}\,\mathrm{TNR},

where πk\pi_{k} is the prior of class kk. When π1\pi_{1} is small, accuracy is dominated by π0 TNR\pi_{0}\,\mathrm{TNR}, and a model can be near-perfect on TNR\mathrm{TNR} while TPR\mathrm{TPR} is collapsed to zero. The metric does not surface the failure.

The metrics that do are

precision  =  TPTP+FP,recall  =  TPTP+FN.\mathrm{precision} \;=\; \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP}}, \qquad \mathrm{recall} \;=\; \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FN}}.

Precision conditions on the predicted positives and asks how clean they are; recall conditions on the real positives and asks how many were caught. Their harmonic mean is F1:

F1  =  2 precision⋅recallprecision+recall.F_{1} \;=\; \frac{2\,\mathrm{precision}\cdot\mathrm{recall}}{\mathrm{precision} + \mathrm{recall}}.

Both precision and recall ignore true negatives entirely — the count that swells when the prior is skewed — so they cannot be inflated by predicting "no" on everything. This is the structural reason PR-style metrics do not lie on imbalanced data the way accuracy does.

MetricUses TN?Skewed-prior safe?What it answers
Accuracyyesno — dominated by majorityoverall hit rate
Precisionnoyesof predicted +, how many are real +
Recallnoyesof real +, how many we caught
F1noyesharmonic mean of P and R
ROC-AUCyesmisleadingly "stable"rank quality vs the diagonal
PR-AUCnoyesprecision achievable at every recall

ROC and PR curves: the same scores, different summaries

A classifier that outputs a continuous score can be thresholded at every possible cut to produce a different confusion matrix. The ROC curve plots FPR\mathrm{FPR} against TPR\mathrm{TPR} as the threshold sweeps from 1 down to 0; the PR curve plots recall\mathrm{recall} against precision\mathrm{precision} over the same sweep.

Both are summaries of one scored classifier. The difference is what they hold fixed.

For ROC, the xx-axis is FPR=FP/(FP+TN)\mathrm{FPR} = \mathrm{FP}/(\mathrm{FP}+\mathrm{TN}). As the prior skews, TN\mathrm{TN} grows and the absolute count of FP\mathrm{FP} needed to reach a given FPR grows too — but the rate stays the same. ROC-AUC is therefore reported as "robust to class imbalance", but that is a property of the ratio, not of the trade-off the practitioner cares about.

For PR, the yy-axis is precision, which by definition does not involve TN\mathrm{TN}. A classifier that lives at precision 0.05 on a 1:99 stream is doing something different from one at precision 0.05 on a 1:1 stream — the same score region means different things — and PR preserves that information.

Mathematically, the two curves are related by the prior. Given a point (FPR,TPR)(\mathrm{FPR}, \mathrm{TPR}) on the ROC curve, the corresponding PR point is

precision  =  π1 TPRπ1 TPR+π0 FPR.\mathrm{precision} \;=\; \frac{\pi_{1}\,\mathrm{TPR}}{\pi_{1}\,\mathrm{TPR} + \pi_{0}\,\mathrm{FPR}}.

A curve that looks fine on ROC (AUC 0.95) can collapse on PR (AUC 0.30) once you read off the actual precision at the recall you'll operate at. Davis and Goadrich showed in 2006 that the two AUCs are monotone in one another — one dominates iff the other does — but the visual quality is dominated by the prior, and the PR plot shows it.

import numpy as np
from sklearn.metrics import precision_recall_curve, roc_curve

def plot_both_curves(y_true: np.ndarray, scores: np.ndarray) -> None:
    """
    Compare ROC and PR for the SAME classifier on the SAME test set.

    Demonstrates that a high ROC-AUC can co-exist with a flat PR curve when
    the prior is skewed.
    """
    fpr, tpr, _ = roc_curve(y_true, scores)
    prec, rec, _ = precision_recall_curve(y_true, scores)

    print(f"ROC-AUC proxy (TPR at FPR=0.1): {np.interp(0.1, fpr, tpr):.3f}")
    print(f"PR-AUC proxy (P at R=0.5):      {np.interp(0.5, rec[::-1], prec[::-1]):.3f}")

The script prints two numbers from the same model. They are not equal, and they are not interchangeable. The second is the one that tells you whether your alert stream is going to be 80% noise or 30% noise at the operating point you care about.

Cost-sensitive learning at training time

There are three places to address imbalance: in the data, in the loss, or at the decision boundary. The data-level fixes (SMOTE etc.) are covered below; here we address the loss.

For logistic regression with weights θ\theta, the per-example log-loss is

ℓi(θ)  =  −[yilog⁡p^i+(1−yi)log⁡(1−p^i)].\ell_i(\theta) \;=\; -\Bigl[ y_i \log \hat{p}_i + (1 - y_i)\log(1 - \hat{p}_i) \Bigr].

The standard empirical risk is 1n∑iℓi\frac{1}{n}\sum_i \ell_i. Under a 1:99 prior, the minority class contributes roughly 1% of the per-example loss regardless of which specific positive example it is, so the gradient is dominated by the negatives. The optimum it converges to is whichever θ\theta makes the negatives well-classified — which is just predicting 0.

Reweighting by the inverse prior class frequency multiplies the minority term by a factor w1>1w_1 > 1:

R(θ)  =  1n ∑i=1n wyi ℓi(θ),w1=n2 n1,  w0=n2 n0.R(\theta) \;=\; \frac{1}{n}\,\sum_{i=1}^{n}\,w_{y_i}\,\ell_i(\theta), \qquad w_1 = \frac{n}{2\,n_1},\; w_0 = \frac{n}{2\,n_0}.

The effective contribution of each class to the gradient is now equal. The model's logits are unchanged in shape, but the decision they map to — the threshold τ\tau that recovers the prior-weighted Bayes rule — is shifted.

The implementation in scikit-learn is the class_weight="balanced" argument. It is not a free improvement: it is the same loss minimised under a different objective, and the resulting classifier will be worse on the original accuracy metric even when it is better on every class-aware metric. The reason is exactly the same as the threshold derivation above — accuracy is a weighted average with weights πk\pi_k, and reweighting the loss changes the metric you implicitly optimise.

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import f1_score, classification_report

def class_weighted_logreg(X_train, y_train, X_test, y_test):
    """
    Train logistic regression with class re-weighting and report F1.

    `class_weight="balanced"` sets w_k = n / (K * n_k), equalising the per-class
    gradient contribution. Expect recall on the minority class to rise sharply
    while overall accuracy falls.
    """
    clf = LogisticRegression(class_weight="balanced", max_iter=2000)
    clf.fit(X_train, y_train)
    y_pred = clf.predict(X_test)
    print(classification_report(y_test, y_pred, digits=3))
    return clf, y_pred

The assumption under which class reweighting is correct: the loss surface is the same shape and only the relative gradient magnitudes are wrong. When that is false — when the minority class is also structurally harder (fewer samples, higher noise, no clean separator) — reweighting simply turns a confidently-bad classifier into a less-confidently-bad one. The only honest cure is more data, not a heavier weight.

Resampling: under-sampling, over-sampling, SMOTE

The third intervention is at the data layer. There are three families.

Random under-sampling drops majority-class examples at random until the class histogram is balanced. It throws away information. For a 1:99 dataset, it discards 98% of the negatives. The remaining sample is statistically small, the gradient is noisy, and the resulting model has high variance. It is fast and works in a notebook demo, and it should not be the production answer.

Random over-sampling duplicates minority examples until the histogram is balanced. It does not throw away information but it does inflate the minority class's effective weight in a way the loss cannot distinguish from class reweighting — and it produces a model that memorises specific positive examples, because each positive shows up multiple times in the same epoch. SMOTE is the named fix for that memorisation.

SMOTE (Synthetic Minority Over-sampling Technique, Chawla et al. 1999) generates new minority examples by interpolating between an existing minority point and one of its kk minority-nearest neighbours:

xnew  =  xi  +  λ⋅(xnn−xi),λ∼Uniform(0,1).x_{\text{new}} \;=\; x_i \;+\; \lambda \cdot (x_{\text{nn}} - x_i), \qquad \lambda \sim \mathrm{Uniform}(0, 1).

The new point lies on the segment between the two originals, so it is not a duplicate and it does extend the convex hull of the minority class. The randomness is over λ\lambda, not over the choice of (xi,xnn)(x_i, x_{\text{nn}}), which is what makes SMOTE a generative procedure rather than a memorising one.

import numpy as np

def smote_one(X_min: np.ndarray, k: int = 5, n_new: int = 1,
              rng: np.random.Generator | None = None) -> np.ndarray:
    """
    Generate `n_new` synthetic minority examples via SMOTE.

    For each new point: pick a minority row x_i, find its k-NN among the
    minority rows, pick one x_nn, and interpolate with random lambda in (0, 1).
    """
    rng = rng or np.random.default_rng(0)
    out = np.empty((n_new, X_min.shape[1]), dtype=X_min.dtype)
    for j in range(n_new):
        i = rng.integers(0, len(X_min))
        diffs = X_min - X_min[i]
        d2 = np.einsum("ij,ij->i", diffs, diffs)
        nn_idx = np.argpartition(d2, k + 1)[1: k + 1]
        nn = rng.choice(nn_idx)
        lam = rng.uniform(0.0, 1.0)
        out[j] = X_min[i] + lam * (X_min[nn] - X_min[i])
    return out

There are two failure modes. First, synthetic noise: if xix_i is itself an outlier or a mislabeled example, SMOTE happily generates new points along its trajectory — polluting the minority class with synthetic garbage. Second, class boundary bleed: when the minority class and the majority class overlap, SMOTE generates points across the boundary, blurring the decision surface. Both are addressed by variants. Borderline-SMOTE generates only from minority points whose kk-NN contain a majority neighbour, which is exactly the boundary zone where new examples help. Tomek links identify pairs of opposite-class neighbours and drop the majority member, sharpening the boundary back. SVMSMOTE uses the SVM support vectors to seed the interpolation, again biasing toward the boundary.

The assumption under which SMOTE is correct: the minority class manifold is smooth, and linear segments between neighbours stay on-manifold. When the manifold is twisted or when the minority cluster is itself fragmented, the interpolations fall off-manifold and the new points are garbage.

Choosing the operating point in production

A trained, thresholded model in production is a decision rule, not a probability estimator. The probability it outputs is the means to the rule, not the rule itself. Picking the rule is a business decision that turns on three numbers the model has no opinion about:

  • C10C_{10}, the cost of a missed positive,
  • C01C_{01}, the cost of a false alarm,
  • the operational capacity (how many alerts per day the team can actually handle).

The Bayes-optimal threshold τ⋆=C01/(C01+C10)\tau^\star = C_{01}/(C_{01}+C_{10}) is the floor: any threshold above τ⋆\tau^\star is provably dominated by a lower one on the cost objective. Capacity is the ceiling: pick the smallest threshold that does not exceed the team's daily budget of false alarms.

The full decision is therefore:

τ  =  min⁡{ τ′  ∣  τ′≥τ⋆  and  FP(τ′)≤capacity }.\tau \;=\; \min\Bigl\{\,\tau' \;\big|\; \tau' \ge \tau^\star \;\text{and}\; \mathrm{FP}(\tau') \le \mathrm{capacity} \,\Bigr\}.

If no such τ′\tau' exists, the model is not yet strong enough to deploy — adding capacity is not a model decision.

import numpy as np

def choose_threshold(y_true: np.ndarray, scores: np.ndarray,
                     cost_fn: float, cost_fp: float,
                     capacity: int) -> tuple[float, dict]:
    """
    Pick the operating threshold that satisfies both Bayes-optimality
    and an operational capacity constraint.
    """
    from sklearn.metrics import roc_curve
    fpr, tpr, thr = roc_curve(y_true, scores)
    # thr has len(fpr)+1 entries; align by trimming
    n = len(fpr)
    fp_counts = fpr * (y_true == 0).sum()
    tau_star = cost_fp / (cost_fp + cost_fn)

    candidates = np.where((thr[:n] <= tau_star) & (fp_counts <= capacity))[0]
    if len(candidates) == 0:
        raise RuntimeError("no threshold satisfies both constraints")
    j = candidates[np.argmin(thr[candidates])]
    return float(thr[j]), {"fpr": float(fpr[j]),
                           "tpr": float(tpr[j]),
                           "tau_star": float(tau_star)}

The function returns the smallest threshold that still beats the Bayes floor and stays within the team's daily false-alarm budget. There is no accuracy number anywhere in the call: accuracy is not the objective.

The failure mode practitioners hit is forgetting the prior. They ship a model whose score distribution is unchanged from the training set, set τ=0.5\tau=0.5 "to be safe", and watch the alert volume collapse to a tenth of what the budget allows. The model is fine. The threshold is wrong. The fix is one number, not a retraining.

Calibration and reliability on imbalanced data

The Bayes threshold derivation assumed the score p^(y=1∣x)\hat{p}(y=1 \mid x) is calibrated — that is, among all points the model scores at value pp, the realised positive rate is approximately pp. On an imbalanced training set, two things go wrong at once.

The first is the prior shift between training and deployment. A model trained on a 1:99 stream learns a score whose mean is π1train=0.01\pi_1^{\mathrm{train}} = 0.01. At deployment, if the prior is 1:1000, the mean score is π1deploy=0.001\pi_1^{\mathrm{deploy}} = 0.001 but the rank order is the same — every threshold that separates positives from negatives in deployment must be re-scaled. The ratio π1deploy/π1train\pi_1^{\mathrm{deploy}} / \pi_1^{\mathrm{train}} is the recalibration factor; it is the prior correction applied uniformly to every score.

The second is the reliability curve. Bin the test points by predicted score; in each bin, plot the realised positive rate against the mean predicted score. On a 1:99 stream, the bins are tiny near the top of the score range — there are not enough positives in the highest bin to estimate the realised rate reliably, and the rightmost point of the reliability curve has wide error bars. Platt scaling and isotonic regression both recalibrate by fitting a one-dimensional mapping g:[0,1]→[0,1]g: [0,1] \to [0,1], but isotonic regression has more capacity (it is piecewise constant) and so overfits on small bins, while Platt is monotone and linear in log-odds and underfits when the miscalibration is non-monotone.

import numpy as np
from sklearn.calibration import calibration_curve

def reliability_table(y_true: np.ndarray, scores: np.ndarray,
                      n_bins: int = 10) -> np.ndarray:
    """
    Bin test points by predicted score and report realised positive rate
    in each bin, together with the bin midpoint.
    """
    edges = np.linspace(0.0, 1.0, n_bins + 1)
    out = np.empty((n_bins, 3))
    for k in range(n_bins):
        mask = (scores >= edges[k]) & (scores < edges[k + 1])
        if mask.sum() == 0:
            out[k] = (0.5 * (edges[k] + edges[k + 1]), np.nan, 0)
        else:
            out[k] = (scores[mask].mean(),
                      y_true[mask].mean(),
                      mask.sum())
    return out

The Brier score BS=1n∑i(yi−p^i)2\mathrm{BS} = \frac{1}{n}\sum_i (y_i - \hat{p}_i)^2 is the mean squared error of the score against the 0/1 label. It is minimised when p^i=Pr⁡(yi=1∣xi)\hat{p}_i = \Pr(y_i = 1 \mid x_i), so on a calibrated classifier it is the squared L2 distance from the Bayes-optimal probability — it is the natural loss for a probability forecaster. On an imbalanced set, Brier is also dominated by the majority class for the same reason accuracy is: a model that predicts p^i≈π1\hat{p}_i \approx \pi_1 everywhere has Brier equal to π1(1−π1)≈π1\pi_1(1-\pi_1) \approx \pi_1, which on 1:99 is 0.00990.0099, and the per-class breakdown BS1=E[(p^−1)2∣y=1]\mathrm{BS}_1 = \mathbb{E}[(\hat{p} - 1)^2 \mid y=1] and BS0=E[p^2∣y=0]\mathrm{BS}_0 = \mathbb{E}[\hat{p}^2 \mid y=0] tells the practitioner where the calibration is failing.

Macro-average, micro-average and the threshold sweep

When the dataset has K>2K > 2 classes or when the class histogram varies across deployment slices, the aggregation choice matters as much as the base metric.

Macro-F1 averages F1 across classes equally:

F1macro  =  1K ∑k=1K F1,k.F_1^{\mathrm{macro}} \;=\; \frac{1}{K}\,\sum_{k=1}^{K}\,F_{1,k}.

Micro-F1 pools the per-class TP, FP, FN counts before computing the ratio:

F1micro  =  2∑kTPk2∑kTPk+∑kFPk+∑kFNk.F_1^{\mathrm{micro}} \;=\; \frac{2\sum_k \mathrm{TP}_k}{2\sum_k \mathrm{TP}_k + \sum_k \mathrm{FP}_k + \sum_k \mathrm{FN}_k}.

Macro and micro are the same on a balanced dataset and very different on a skewed one. Micro-F1 is dominated by the majority class — when n1≫n0n_1 \gg n_0, the F1 reported is essentially the F1 of the majority class. Macro-F1 is dominated by the worst class — when one class is rare and difficult, macro-F1 falls sharply regardless of how well the others do. The 201-level rule is to report macro-F1 when the question is "does the model work for every class?" and micro-F1 (or weighted F1) when the question is "does the model work for the average example?".

A common mistake is to compute F1 with the default 0.5 threshold and report it as the model's quality. F1 at τ=0.5\tau = 0.5 is one point on the PR curve; the best F1 is the maximum over τ\tau, and best-F1 ignores the operating point that production actually needs. The threshold sweep that gives best-F1 is the same sweep that defines the PR curve, and the area under the PR curve is the average precision — a one-number summary that does not depend on which operating point you happened to choose for the report.

Worked example: a 1:99 stream, end to end

The following walks the full pipeline on a synthetic 1:99 stream and shows the failure of every 101-level instinct.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (precision_recall_curve, roc_curve,
                             f1_score, average_precision_score,
                             roc_auc_score)

rng = np.random.default_rng(0)
n, p = 50_000, 20
X = rng.normal(size=(n, p))
# Positives are a small cluster in feature space
y = (X[:, 0] + 0.5 * X[:, 1] - 0.3 * X[:, 2] > 1.2).astype(int)
# Force a 1:99 prior
neg, pos = np.where(y == 0)[0], np.where(y == 1)[0]
keep_neg = rng.choice(neg, size=99 * len(pos), replace=False)
keep = np.concatenate([pos, keep_neg])
rng.shuffle(keep)
X, y = X[keep], y[keep]

# Train / test split, then fit baseline and reweighted logreg
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25,
                                          stratify=y, random_state=0)

m_default = LogisticRegression(max_iter=2000).fit(X_tr, y_tr)
m_balanced = LogisticRegression(class_weight="balanced",
                                max_iter=2000).fit(X_tr, y_tr)

for name, m in [("default", m_default), ("balanced", m_balanced)]:
    s = m.predict_proba(X_te)[:, 1]
    auc_roc = roc_auc_score(y_te, s)
    auc_pr = average_precision_score(y_te, s)
    f1_05 = f1_score(y_te, (s > 0.5).astype(int))
    # Best F1 by sweeping the threshold
    pr, rc, thr = precision_recall_curve(y_te, s)
    f1_curve = 2 * pr * rc / np.maximum(pr + rc, 1e-12)
    f1_best = f1_curve[:-1].max()
    print(f"{name:>9s}: ROC-AUC={auc_roc:.3f}  PR-AUC={auc_pr:.3f}  "
          f"[email protected]={f1_05:.3f}  best-F1={f1_best:.3f}")

The printout is the lesson's whole argument in four numbers per model. The default logreg has high ROC-AUC (the rank order is fine), low PR-AUC (the precision at every useful recall is poor), middling [email protected] (the threshold is wrong), and a best-F1 that requires a τ≪0.5\tau \ll 0.5. The balanced logreg has lower ROC-AUC (its rank order is worse because the loss is different) but higher PR-AUC and a best-F1 that is achievable at a defensible threshold.

The 201-level mistake the printout guards against: shipping the default logreg because its ROC-AUC is highest. The 101-level instinct is to look at one number; the 201-level practice is to look at the curve and the operating point together, with cost and capacity constraints explicit. The threshold is not "0.5 by default" — it is "τ⋆\tau^\star shifted by capacity", and the model's score is a means to that threshold, not an end in itself.

Key Takeaways

  • The Bayes-optimal threshold under unequal costs is τ⋆=C01/(C01+C10)\tau^\star = C_{01}/(C_{01}+C_{10}); the textbook 0.5 is the special case C01=C10C_{01}=C_{10}, not the general rule.
  • Accuracy is π1 TPR+π0 TNR\pi_1\,\mathrm{TPR} + \pi_0\,\mathrm{TNR} and is dominated by the majority class; F1, precision and recall do not involve TN\mathrm{TN} and cannot be inflated by always predicting the majority.
  • ROC-AUC and PR-AUC are monotone in one another, but on skewed data the PR curve is what reveals the precision/recall trade-off at the operating point you will actually use.
  • Class reweighting equalises gradient contributions from each class but does not invent information; SMOTE extends the minority convex hull but does not repair mislabels or class-overlap noise.
  • The shipping threshold is the smallest value above τ⋆\tau^\star that respects the team's false-alarm capacity; if no such threshold exists, the model is not ready, not "tuned too conservative".
  • Calibration, Brier score, macro-F1 and micro-F1 each expose a different aspect of imbalanced performance that a single ROC-AUC number silently hides.

Check your understanding

8 questions · 80% to complete the lesson

1 / 8

7 correct to pass

A fraud model scores every transaction with p^(fraud=1∣x)\hat{p}(\text{fraud}=1\mid x). The cost of missing fraud is $1,000 and the cost of a false-positive block is $50. According to the Bayes-optimal rule, at what score should the model alert?

0 of 8 answered

Pick a lesson to start the audio.