07

Feature Engineering & Pipelines

Cantonese podcast title: 特徵工程與管道

Learning Objectives

  1. Define target leakage and identify the three places it typically enters a feature pipeline (training statistics, target-derived encodings, time-aware splits).
  2. Build a ColumnTransformer that applies different transformations to numeric and categorical columns without ever computing a statistic on the validation or test data.
  3. Derive target encoding (mean encoding) from a hierarchical Bayesian perspective and explain why naive mean encoding leaks the target.
  4. Diagnose whether a fitted pipeline is leaky by inspecting what was learned on which data, using cross-validation behaviour as the evidence.
  5. Choose between imputation, scaling, encoding, binning, and interaction features for a given dataset, and justify each choice in terms of the model class that will consume them.
Feature Engineering & Pipelines — visual guide
Pipeline and ColumnTransformer: leakage-free composition Pipeline + ColumnTransformer: fit on train, transform on test the fitted pipeline is the model; never refit on test, never refit on validation Raw columns age income tenure_mo city plan_type device_cls city has 48000 unique values target = churn ColumnTransformer numeric pipe impute(median) -> StandardScaler fit on TRAIN only categorical pipe impute(most_freq) -> OneHotEncoder for city: out-of-fold target encoding + smoothing remainder='drop' -> typo in a column name surfaces as a missing feature, not a silent bug Estimator LogisticRegression or gradient boosting sees only the transformed matrix Prediction y_hat on held-out X Three leak paths: (1) statistics across train/test (2) target-derived features using current row (3) temporal leaks using future info

Assumes you know from ML-101

This lesson builds on ML-101 · Lesson 2 (Data & Features), ML-101 · Lesson 4 (Logistic Regression), and ML-101 · Lesson 9 (Model Evaluation). You should already understand the difference between a feature and a label, know how to one-hot encode a categorical column and standardise a numeric column, and remember that the train/test split exists to measure generalization. We will treat those as starting points, not as topics to re-explain.

You should also remember from ML-101 · Lesson 9 that evaluation metrics are computed on a held-out set, and that the held-out set must be untouched during model fitting. This lesson is largely about how to keep it untouched when the feature engineering itself uses information that would leak if applied carelessly.

Learning Objectives

  1. Define target leakage and identify the three places it typically enters a feature pipeline (training statistics, target-derived encodings, time-aware splits).
  2. Build a ColumnTransformer that applies different transformations to numeric and categorical columns without ever computing a statistic on the validation or test data.
  3. Derive target encoding (mean encoding) from a hierarchical Bayesian perspective and explain why naive mean encoding leaks the target.
  4. Diagnose whether a fitted pipeline is leaky by inspecting what was learned on which data, using cross-validation behaviour as the evidence.
  5. Choose between imputation, scaling, encoding, binning, and interaction features for a given dataset, and justify each choice in terms of the model class that will consume them.

What leakage is, exactly

Target leakage happens when information that would not be available at prediction time influences the training process. The trained model then appears to perform well on the held-out set — because the held-out set was, in some sense, partially memorized — and performs poorly in production, where the leaked information is genuinely absent.

A leak is not the same as overfitting. An overfit model fails because it has too much capacity relative to its data; a leaky model fails because it has been given access to the answers it is supposed to predict. The two have similar symptoms (a held-out metric that does not transfer to production) but different causes and different remedies. Overfitting is fixed by regularisation, more data, or a simpler model. Leakage is fixed by changing the pipeline so the leaked information never enters the model.

The cleanest diagnostic is to ask: "if I had a fresh example tomorrow, would I be able to compute this feature for it using only the data I would have tomorrow?" If the answer is no, the feature is leaky. Some leaks are obvious (the label is included as a feature), but most are subtle (the count of training examples in the same category as the test example is computed over both splits), and the subtle ones are the dangerous ones because they survive code review.

The three leak paths

Leakage enters through one of three doors. Knowing the doors is what lets you recognise a leak when you see one.

1. Statistics computed across the train/test boundary. If you standardise a column using the mean and standard deviation of the whole dataset (train + test), the test examples have contributed to those statistics, and any model fit on the standardised training data has implicitly seen the test distribution. The fix is mechanical: fit the standardiser on the training data only, then apply it (transform, not fit) to the validation and test data. The same rule applies to imputation, scaling, PCA, target encoding, vocabulary construction, and any other operation that aggregates over rows.

2. Target-derived features that include the current row. A "category mean target" feature for a category that contains the current row is a leak if that row's own target contributed to the mean. The standard remedy is to use out-of-fold target encoding: estimate the mean target for category cc using only the rows not in the current fold, then assign that estimate to every row in the fold. Cross-validated target encoding is the standard implementation.

3. Temporal leaks in time-series data. A feature that uses future information at the time the prediction is made is a leak even if both train and test rows are in the same dataset. The remedy is walk-forward validation (covered in a later lesson) and rolling statistics that use only past values. A common subtle case: computing "user's average purchase over the past 30 days" using the entire history including days that are part of the test period.

The third path is the most painful because it requires understanding the data-generating process, not just the data. The first two can be caught by disciplined pipeline construction. The third requires you to ask, for every feature, "when would this value be available in production?"

Why pipelines are not optional

The scikit-learn Pipeline class chains a sequence of transformer/estimator pairs so that the entire sequence can be fit and predict-ed as one object. The critical property is that when the pipeline is fit on training data, every transformer in the chain is fit on the training data only. When the pipeline is predict-ed on new data, every transformer is transform-ed (not refit) on the new data.

This means a pipeline is not just convenience — it is the only way to guarantee that test-time features are computed the same way as training-time features, using statistics that were learned only from training data. If you compute features in a notebook by hand, you cannot reproduce the exact computation on a new example without explicitly carrying around the fitted transformer objects. If you put the same computation inside a pipeline, the fitted pipeline is the artefact you save and ship.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

# A pipeline is a chain of (name, transformer-or-estimator) pairs.
# When you call pipeline.fit(X_train, y_train), the StandardScaler is
# fit on X_train only, and the LogisticRegression sees only the scaled
# training data. When you call pipeline.predict(X_test), the StandardScaler
# is *applied* to X_test using the mean and sd learned on X_train, never
# refit on X_test.
pipe = Pipeline([
    ("scale", StandardScaler()),
    ("clf", LogisticRegression(max_iter=1000)),
])
pipe.fit(X_train, y_train)
preds_test = pipe.predict(X_test)

The mental model is: Pipeline.fit and Pipeline.predict propagate fit/transform through the chain in order. A transformer that is not in the chain cannot leak by definition, because the model never sees its output.

ColumnTransformer: composition without leakage

Real datasets have heterogeneous columns: numeric, categorical, ordinal, text, and possibly missing-everywhere. A single StandardScaler cannot handle a categorical column, and a single OneHotEncoder cannot handle a numeric one. The scikit-learn ColumnTransformer lets you specify different transformations for different subsets of columns and concatenate the results.

The mechanical rule for using ColumnTransformer without leakage is the same as for Pipeline: the ColumnTransformer.fit happens on training data only. Each underlying transformer is fit on its own subset of the training data, never on the validation or test data. The fitted ColumnTransformer is then applied to validation and test as a single transform call.

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline

# Numeric columns get median imputation + standardisation.
numeric_pipe = Pipeline([
    ("impute", SimpleImputer(strategy="median")),
    ("scale", StandardScaler()),
])

# Categorical columns get most-frequent imputation + one-hot encoding.
categorical_pipe = Pipeline([
    ("impute", SimpleImputer(strategy="most_frequent")),
    ("onehot", OneHotEncoder(handle_unknown="ignore")),
])

# ColumnTransformer composes the two pipelines.
# When you call .fit on this object, each sub-pipeline is fit on its own
# column subset of the training data only. When you call .transform on
# new data, each sub-pipeline is applied (not refit) to its own subset.
preprocess = ColumnTransformer(
    transformers=[
        ("num", numeric_pipe, ["age", "income", "tenure_months"]),
        ("cat", categorical_pipe, ["city", "plan_type", "device_class"]),
    ],
    remainder="drop",  # any column not listed is dropped, not silently kept
)

The remainder="drop" default is important: a column that is not explicitly listed in transformers is dropped, which means a typo in a column name produces a visible error (the column disappears from the model) rather than a silent bug (the column is fed in raw form and the model sees a mixture of scaled and unscaled values).

The fitted ColumnTransformer produces a numpy array or sparse matrix. The columns are no longer named; the only way to recover which input feature produced which output column is via preprocess.get_feature_names_out(). This is a real cost for production code: pipelines that need to be inspected by humans must retain the feature-name mapping somewhere.

Target encoding and its leakage problem

A categorical column with high cardinality (city names, product IDs, browser user-agent strings) is impractical to one-hot encode: a single column with 50,000 unique values would add 50,000 sparse columns, most of them zero. Target encoding replaces the categorical value with the mean of the target variable for that value.

Concretely, for a binary target y∈{0,1}y \in \{0, 1\} and a categorical column cc taking values in C\mathcal{C}, the target encoding is

enc(c)  =  1∣{i:ci=c}∣∑i:ci=cyi\text{enc}(c) \;=\; \frac{1}{|\{i : c_i = c\}|} \sum_{i : c_i = c} y_i

This is just the conditional mean E[y∣c]E[y \mid c], and it is a reasonable summary of how the target depends on the category. A linear or tree model can use it directly.

The leak is immediate: the encoding for a category that contains the current row uses that row's own target. A classifier fit on this encoding sees, for each row, a feature value that already contains the answer. The held-out evaluation looks excellent because the classifier has effectively memorized the labels through the encoding.

The fix is out-of-fold target encoding: split the training data into KK folds. For each fold, estimate the encoding using only the other K−1K-1 folds, then assign that encoding to the current fold. Every row's encoded feature is computed from data that does not include its own target, but the estimate is still made on the full training distribution (because the other folds together cover all the data).

import numpy as np
from sklearn.model_selection import KFold

def out_of_fold_target_encode(c_train, y_train, c_val=None, n_splits=5, prior=None):
    """Compute out-of-fold target encoding for c_train, and full-data encoding for c_val.

    For each row in c_train, the encoding is computed from the other n_splits-1 folds.
    For c_val (held-out data), the encoding is computed from the entire training set.
    The `prior` argument lets the caller smooth the estimate toward the global mean,
    which is essential for low-count categories.
    """
    if prior is None:
        prior = float(np.mean(y_train))

    enc_train = np.zeros(len(c_train), dtype=float)
    kf = KFold(n_splits=n_splits, shuffle=True, random_state=0)
    for tr_idx, va_idx in kf.split(c_train):
        # Compute per-category mean on the training fold.
        df = (
            pd.DataFrame({"c": c_train[tr_idx], "y": y_train[tr_idx]})
            .groupby("c")["y"].agg(["mean", "count"])
        )
        # Look up each row's encoding from the training fold only.
        enc_train[va_idx] = (
            c_train[va_idx]
            .map(lambda v: df.loc[v, "mean"] if v in df.index else prior)
        )

    # For the held-out set, refit on the *entire* training data.
    df_full = (
        pd.DataFrame({"c": c_train, "y": y_train})
        .groupby("c")["y"].agg(["mean", "count"])
    )
    enc_val = (
        c_val.map(lambda v: df_full.loc[v, "mean"] if v in df_full.index else prior)
        if c_val is not None
        else None
    )
    return enc_train, enc_val

The smoothing toward prior is the next refinement, and it addresses a related but distinct problem: a category with only one or two observations has an extremely noisy encoding. Without smoothing, a category with one positive example and zero negatives gets encoded as 1.01.0, which the classifier treats as a strong signal. Bayesian smoothing toward the global mean yˉ\bar{y} shrinks these estimates:

encsmoothed(c)  =  nc⋅yˉc+m⋅yˉnc+m\text{enc}_{\text{smoothed}}(c) \;=\; \frac{n_c \cdot \bar{y}_c + m \cdot \bar{y}}{n_c + m}

where ncn_c is the count of category cc, yˉc\bar{y}_c is the empirical mean, and mm is a smoothing hyperparameter (often m=10m = 10 or m=30m = 30). This is the hierarchical Bayesian target encoding, derived as the posterior mean of a Beta-Binomial conjugate model.

Hierarchical Bayesian target encoding

The Bayesian derivation: model y∣c∼Bernoulli(θc)y \mid c \sim \text{Bernoulli}(\theta_c) independently within each category, and put a prior θc∼Beta(α,β)\theta_c \sim \text{Beta}(\alpha, \beta) on the per-category rates, with hyperparameters α,β\alpha, \beta shared across all categories. The posterior is

θc∣y  ∼  Beta(α+∑i:ci=cyi,  β+∑i:ci=c(1−yi))\theta_c \mid y \;\sim\; \text{Beta}\Bigl(\alpha + \sum_{i : c_i = c} y_i,\; \beta + \sum_{i : c_i = c} (1 - y_i)\Bigr)

The posterior mean is

E[θc∣y]  =  α+ncyˉcα+β+ncE[\theta_c \mid y] \;=\; \frac{\alpha + n_c \bar{y}_c}{\alpha + \beta + n_c}

Setting α=myˉ\alpha = m \bar{y} and β=m(1−yˉ)\beta = m(1 - \bar{y}) gives the smoothed encoding above. The shrinkage factor m/(m+nc)m / (m + n_c) controls how much a low-count category's estimate is pulled toward the prior; categories with thousands of observations are barely affected.

This is the same smoothing logic as the empirical Bayes estimator for the Poisson rate, and it shows up in language modelling (Laplace smoothing, Kneser-Ney), recommender systems (shrink toward the global click rate), and survival analysis (shrink toward the median lifetime). The lesson is general: any per-category rate is at risk of being noisy for low counts, and the right shrinkage comes from a hierarchical Bayesian model with a shared prior.

Time-aware features and rolling statistics

For time-series or panel data, "the user's average spend over the past 30 days" is a common feature, and it is a common leak: if the rolling window includes the day being predicted, the feature contains the answer. The fix is the strict rolling statistic: at time tt, the feature uses only information from times strictly before tt.

import pandas as pd

# A leak-free rolling mean: at each row, average over the previous k-1 rows.
# The shift(1) excludes the current row from the rolling window.
df["rolling_30d_avg"] = (
    df.groupby("user_id")["spend"]
      .shift(1)                        # exclude current row, prevent leak
      .rolling(window=30, min_periods=1)
      .mean()
      .reset_index(level=0, drop=True)
)

The shift(1) is the load-bearing line. Without it, the rolling window includes the current row and the feature is leaky. With it, the feature at time tt is computed from information that would have been available at the moment prediction tt is made. This is the single-line rule for time-series feature engineering: never use the current row's label or value in any feature that will be used to predict the current row's label.

A second, less obvious leak in time-series features is the use of future observations in the rolling window. Computing "user's average spend over the next 7 days" is a feature that no production system can produce. The walk-forward validation in a later lesson is the diagnostic that catches this: if a model trained with future-looking features performs well in standard k-fold cross-validation but badly under walk-forward validation, the features are using future information.

Scaling, and the one transform that is not optional

Two of the "optional" preprocessing steps are in fact load-bearing, and the reason is scale-dependent optimisation rather than taste.

Standardisation for gradient-based models. A linear model's predictions are invariant to affine rescaling of the features — rescale a column, the fitted coefficient divides by the scale, the fitted value is identical. The loss surface is not invariant. For unstandardised features with wildly different units, the Hessian's diagonal spans orders of magnitude, the surface is a narrow valley, and the gradient is nearly orthogonal to the valley floor. Plain gradient descent then zig-zags along the floor almost without progress. Adam's per-coordinate rescaling makes it much less sensitive, which is why the problem survives in deep learning — but any model trained by SGD, by coordinate descent with a fixed learning rate, or by L-BFGS is still affected. For k-means the scale dependence is total rather than partial: the algorithm depends only on Euclidean distance, so a feature in millimetres dominates one in kilometres.

Log transforms for heavy tails. A feature whose distribution is right-skewed — income, session count, order value — has a mean that is not representative of a typical value, and a scale-sensitive distance metric will be dominated by a small number of enormous values. The log compresses the tail while preserving order, and it turns the relationship with the target into something closer to linear. It is also the transform that makes a multiplicative relationship additive, so it is the right first move whenever the residuals fan out with increasing magnitude — that is a log-normal error, not a linear one.

SituationTransformWhySymptom if skipped
Gradient-based model, mixed unitsStandardise (z-score)Conditions the HessianZig-zag SGD; long plateau before loss falls
Tree model, mixed unitsNoneSplits are invariant to monotone rescaling of one featureNo effect — adding it only costs a fit
k-means / k-NN / PCAStandardiseDistance-based; scale is not ignorableClustering by the unit, not the signal
Skewed count or money featurelog⁡(1+x)\log(1 + x)Compresses the right tailResiduals fan out; R² looks fine, MAE does not
Heavy-tailed target, constant variancelog⁡y\log yMakes the error model additiveNon-constant error variance; biased predictions

The middle row is the one people get wrong in both directions. Standardising before a gradient-boosted tree is a common cargo-cult habit copied from linear-model tutorials: it adds a fitted step, a possible source of train/test skew, and buys literally nothing, because a tree's split on income > 50000 means the same thing whatever units it is in. It is harmless but it is not free, and "harmless but not free" is a good reason to know why.

Encoding: choosing by cardinality and by what the model is

The encoding decision is a function of the cardinality of the categorical and the model family, and getting it wrong is one of the largest single sources of wasted capacity in tabular modelling.

One-hot is correct for low cardinality and for any model that needs to see interactions. Its problem is linear growth: 200 categories means 200 columns, and with LL layers of one-hot input the first weight matrix grows by 200L200L parameters. For a model with an embedding table, one-hot on 10,000 categories is simply not an option.

Ordinal encoding assigns an integer per category. It is only correct when the categories are genuinely ordered (low/medium/high risk). Applied to unordered nominal categories it asserts a false ordering, and a linear model will happily use it as though "category 7 is between 6 and 8". This is a silent, hard-to-detect bug: performance degrades rather than crashing.

Target encoding is the workhorse for high-cardinality nominals, with the leakage caveats from the previous section. The key insight is that the encoding itself is a model fit, so it belongs inside the cross-validation fold like any other estimator.

Embeddings learn a vector per category from scratch. They only pay off when there is enough data per category and some categories share structure — with 20 observations for each of 500 categories, a learned embedding is pure overfitting, and a target encoding with smoothing is the better bet. The decision boundary is roughly: embeddings below about 20 observations per category, target encoding above it, and a learned embedding tuned with early stopping when near the boundary.

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer

# The imputer's indicators matter: a missing value is often informative, and
# SimpleImputer(add_indicator=True) preserves that as its own column rather than
# collapsing "missing" and "measured zero" into the same code.
numeric = Pipeline([
    ("impute", SimpleImputer(strategy="median", add_indicator=True)),
    ("scale",  StandardScaler()),
])
categorical = Pipeline([
    ("impute", SimpleImputer(strategy="most_frequent")),
    ("onehot", OneHotEncoder(handle_unknown="infrequent_if_exist", min_frequency=10)),
])

pre = ColumnTransformer(
    [("num", numeric, num_cols), ("cat", categorical, cat_cols)],
    remainder="drop",      # explicit: unknown columns are dropped, not silently passed through
    verbose_feature_names_out=False,
)

Three details in that snippet are load-bearing and all three are routinely omitted. handle_unknown="infrequent_if_exist" with a min_frequency is what stops an unseen test category from raising; the default "error" is a production outage waiting for the first new category. add_indicator=True keeps missingness visible when it is informative. And remainder="drop" is an explicit decision — the default remainder="passthrough" will happily feed a string column to a numeric transform the moment someone adds a column to the dataframe upstream.

The leakage audit: how to actually find it

Leakage is the highest-leverage defect in a tabular project, and it is also the one most often shipped. The reason is structural: leakage produces a good score. There is no failing test, no exception, no crash — the metric improves, and improvement is what you were told to maximise. So it has to be hunted deliberately.

Audit 1 — feature availability. For every feature, write down the timestamp at which it was computed in production. If that is later than the moment the prediction is made, the feature is leaky regardless of how well it scores. This is a table, not a model, and it catches more leaks than anything else.

Audit 2 — single-feature AUC. Fit a model on one feature at a time. A feature with near-perfect AUC is not a strong feature, it is a leak — legitimate features are usually individually mediocre. This is the single highest-yield automated check.

Audit 3 — permutation importance against a shuffled-label control. Fit the full model, then fit it again on shuffled labels. If a feature is important in the real fit but also important in the shuffled fit, the pipeline is leaking through it, because the only way a feature carries signal under shuffled labels is a leak.

Audit 4 — the sanity split. Split by time, by entity, or by group rather than randomly, and check the score still holds. A model that collapses under a group-aware split was memorising group identity, which random splitting actively encourages.

import numpy as np
from sklearn.model_selection import cross_val_score
from sklearn.tree import DecisionTreeClassifier

def single_feature_audit(X, y, features, cv=5, random_state=0):
    """Audit 2. Flags features that are individually too good.

    A legitimate predictive feature is usually mediocre on its own; a leaky one is
    near-perfect. The threshold here is deliberately conservative — it is a review
    trigger, not an automatic removal, because some real problems do have a
    genuinely dominant feature.
    """
    flagged = []
    for f in features:
        auc = cross_val_score(
            DecisionTreeClassifier(max_depth=3, random_state=random_state),
            X[[f]], y, cv=cv, scoring="roc_auc",
        ).mean()
        if auc > 0.90 or auc < 0.10:
            flagged.append((f, float(auc)))
    return flagged

def shuffled_label_control(pipeline, X, y, cv=5, random_state=0):
    """Audit 3. The control arm: the same pipeline on permuted labels must score ~0.5.

    Any pipeline scoring materially above chance here is reading something other than
    the label — almost always a statistic computed across the train/test boundary.
    """
    rng = np.random.default_rng(random_state)
    y_shuf = rng.permutation(y)          # permute LABELS, not rows
    return cross_val_score(pipeline, X, y_shuf, cv=cv, scoring="roc_auc").mean()

The shuffled-label control deserves emphasis because it is the one most projects skip and it is nearly free. It is the empirical answer to "could this pipeline see the answer?" — if shuffling the labels does not destroy the score, the pipeline was never using the labels.

Key Takeaways

  • Target leakage happens when information that would not be available at prediction time influences the training process. The three leak paths are statistics across the train/test boundary, target-derived features that include the current row, and temporal leaks that use future information.
  • Pipeline and ColumnTransformer are not just convenience; they are the only way to guarantee that test-time features are computed the same way as training-time features, using statistics learned only from training data.
  • Target encoding replaces high-cardinality categoricals with the conditional mean of the target, but the naive version leaks. Out-of-fold estimation plus hierarchical Bayesian smoothing is the standard remedy.
  • Time-series features must exclude both the current row and any future observations. The shift(1) in a rolling statistic is the load-bearing operation.
  • Scaling is load-bearing exactly where the model is scale-sensitive: gradient descent, distance-based methods. It is a no-op for tree splits, so applying it there costs a fitted step and buys nothing.
  • The encoding choice is a joint function of cardinality and model family. Ordinal encoding of nominal categories asserts a false ordering, and embeddings below roughly 20 observations per category are pure overfitting.
  • Leakage ships because it improves the metric. Find it with a feature-availability table, a single-feature AUC audit, a shuffled-label control, and a group-aware split — none of which require a new framework.
  • A fitted pipeline is the artefact you ship. The scaler means, the encoder vocabularies, the imputation medians, and the column ordering are all part of the model and must be serialised and versioned with the trained weights.

Check your understanding

6 questions · 80% to complete the lesson

1 / 6

5 correct to pass

Which of the following is the clearest example of target leakage in a feature pipeline?

0 of 6 answered

Pick a lesson to start the audio.