Assumes you know from ML-101
This lesson assumes you have already worked through ML-101 · Lesson 11 (Gradient Descent & Optimization) and ML-101 · Lesson 14 (Neural Networks & Backprop). You should be comfortable with the basic SGD update , with the fact that the gradient points in the direction of steepest ascent, and with the backprop chain rule that produces for every parameter in a feed-forward network. We will not re-derive any of that here; instead, we will start from the assumption that gradients exist and ask why a naive SGD step is a bad idea and what each modern optimizer does about it.
You should also remember from ML-101 · Lesson 10 (Overfitting, Bias & Variance) that a model with too much capacity relative to its data is the standard reason an optimizer appears to "work" on the training set but not on the held-out set. The optimizer stories in this lesson are partly stories about how to fit fast and partly stories about what you are implicitly fitting to — both matter.
Learning Objectives
- Derive the Adam update rule from its components (first and second moment estimates with bias correction) and explain why each component is necessary rather than decorative.
- Distinguish coupled weight decay (an L2 penalty added to the loss) from decoupled weight decay (AdamW) and show the algebraic step that makes them differ for adaptive optimizers.
- Decide when momentum, RMSProp, Adam, or AdamW is the right default for a given problem, and justify the decision in terms of gradient geometry and regularization pressure.
- Diagnose a training loss curve that diverges, plateaus, or oscillates by reading off which optimizer hyperparameter is mis-sized.
- Implement a minimal AdamW loop in NumPy and reproduce the bias-correction math without consulting a library.
Why plain SGD is not enough
SGD with a fixed learning rate updates each parameter using the local gradient :
This single rule is mathematically correct: it follows the negative gradient, which by definition is the direction of steepest descent. The problem is not with the direction but with the magnitude and the conditioning. Suppose two parameters have gradients with very different scales — for instance, a weight on a 0-to-1 feature and a weight on a 0-to-100000 feature, or two parameters whose loss surface is elongated along one axis. A single that is small enough to be stable along the steep axis is too small to make progress along the shallow axis, so the iterates zigzag slowly along a valley floor. This is the ill-conditioning failure mode of plain SGD, and it is the reason every optimizer we will discuss either rescales the step per-coordinate or carries state that smooths the path.
There is a second, more subtle failure mode: stochastic noise. In ML-101 · Lesson 11 we saw that minibatch gradients are unbiased estimates of the full-batch gradient, but they have a non-zero variance that scales inversely with the batch size. With a small batch and a fixed , the iterates bounce around the true descent direction; with a large batch the noise shrinks but the per-step cost grows. SGD does nothing to use the noise structure — every minibatch is treated as an independent perturbation, and the optimizer does not accumulate information across steps. That is the door that adaptive methods open: by remembering past gradients, the optimizer can trade variance for accuracy along axes where the signal is consistent and step cautiously where it is not.
A third reason matters for deep nets specifically. The loss surface is not convex, so the iterates must navigate around saddles and shallow basins. SGD can get trapped in sharp basins that generalize poorly (this is sometimes called "implicit regularization by noise"); a momentum term helps the iterate carry through narrow valleys without being killed by the local curvature. This is the door that momentum opens.
Momentum: smoothing the path
Polyak (heavy-ball) momentum keeps an exponentially weighted moving average of past gradients and uses that as the update direction. The state is a velocity vector that accumulates gradients with decay :
The parameter update is then
The default means is approximately the average of the last ten gradients, with geometrically decaying weights. The dynamics are the discrete analogue of a damped second-order ODE — the iterate picks up speed along consistent directions and damps out oscillating ones. The result, on a valley-shaped loss surface, is that the iterate no longer zigzags between the walls: it accumulates velocity down the valley axis and slows naturally at the bottom.
It is worth being precise about what momentum does. It does not change the expected direction of the update — if the gradients are unbiased, is still proportional to the true gradient. It changes the variance of the update direction and the effective step size along consistent directions. That distinction is why momentum and learning-rate decay interact non-trivially: as training proceeds, the iterates reach a regime where the gradient is small but consistent, and the accumulated velocity produces a longer effective than the nominal one.
Nesterov momentum re-evaluates the gradient at the "look-ahead" point rather than at the current point. The update rule is
Geometrically, Nesterov momentum looks ahead before applying velocity, which gives a small but consistent improvement on convex problems and a slightly better-behaved trajectory on non-convex ones. The difference is rarely decisive in practice; the choice of matters more than the choice between Polyak and Nesterov.
To see momentum in action, consider the recurrence unrolled for a constant gradient direction (the simplest non-trivial case). The velocity is a geometric sum:
After many steps, , the same as SGD. But the transient behaviour differs: a heavy-ball iterate that starts at rest will reach a steady-state velocity in roughly steps, and the effective step size along consistent directions is approximately — for , this is . This is why practitioners report that "Adam with behaves like SGD with ": the momentum is doing most of the work.
def sgd_momentum_step(params, grads, velocity, lr=1e-2, beta=0.9):
"""Polyak heavy-ball momentum: v accumulates gradients with decay beta.
This is the simplest possible adaptive method. The state is just the
velocity vector; the same value is updated for every parameter, with
no per-coordinate rescaling. Watch for velocity overshoot when beta
is close to 1: the iterates can accelerate past a minimum before
the gradient has time to slow them down.
"""
for i, (p, g) in enumerate(zip(params, grads)):
velocity[i] = beta * velocity[i] + g
p -= lr * velocity[i]
return params, velocity
The default for Adam and for SGD-with-momentum in some libraries reflect the fact that Adam has the second moment to lean on for stability, while SGD relies more heavily on momentum. The two are not directly comparable: Adam's momentum is on a rescaled gradient (after dividing by ), while SGD's momentum is on the raw gradient. Setting close to 1.0 in either method gives a longer effective averaging window but at the cost of slower response to changes in the gradient direction. For non-stationary objectives (e.g. reinforcement learning, or training with curriculum learning), a smaller is often better because the optimiser can adapt to the changing loss surface more quickly.
A useful diagnostic for momentum tuning is to plot the running average of the cosine similarity between consecutive gradients. On a smooth loss surface, the similarity is close to 1.0 and momentum helps; on a chaotic surface, the similarity oscillates around 0 and momentum can hurt because the velocity vector is the sum of nearly orthogonal updates and ends up close to zero. If you observe this pattern, the fix is either a smaller or a switch to Adam, whose second moment handles the chaotic-gradient regime better than raw momentum.
Per-coordinate scaling: AdaGrad and RMSProp
Momentum smooths the direction of the update but uses the same learning rate for every coordinate. The next class of methods scales the step per-coordinate using a running estimate of the gradient's second moment. AdaGrad accumulates the sum of squared gradients:
Here is a tiny constant (typically ) that prevents division by zero. The intuition: coordinates that have received consistently large gradients have a large , so the per-step effective learning rate shrinks for them. Coordinates that have received only small gradients keep a larger effective step. This is exactly the per-coordinate rescaling that fixes the ill-conditioning failure of plain SGD.
The problem with AdaGrad is that grows monotonically — it never shrinks — so the effective learning rate decays to zero over the course of training and the optimizer eventually stops moving. This is fine for convex problems where you want aggressive early progress and natural slow-down, but on deep nets it kills long-term learning. RMSProp fixes this by using an exponentially weighted moving average of rather than a cumulative sum:
with typically 0.99. This forgets old gradients and lets the effective learning rate recover if a coordinate starts to receive larger gradients again.
The geometric interpretation of the AdaGrad/RMSProp update is illuminating. The effective per-step movement is , which is approximately a unit-variance coordinate system — coordinates with consistently large gradients have small unit-variance updates, and coordinates with consistently small gradients have large unit-variance updates. This is not the same as normalising the gradient to unit length; the normalisation is per-coordinate, not for the gradient as a whole. The resulting effective Hessian that the optimizer sees is approximately the identity, which is precisely the conditioning fix SGD needed.
A subtle property of this update is that the effective learning rate adapts to the gradient history of each coordinate, not to the local curvature. For convex problems these are the same up to a factor, but on non-convex surfaces the two diverge — AdaGrad and RMSProp can be fooled by a coordinate that recently had a single large gradient spike, which permanently shrinks the effective learning rate for that coordinate (AdaGrad) or for many steps (RMSProp). This is part of why these methods sometimes fail on recurrent networks, where gradient magnitudes are bursty.
The choice of in the per-coordinate scaling has a clear interpretation: it is the maximum effective learning rate any coordinate can receive. When , the effective rate is approximately . With the default , this would be , which is huge. The default is small enough that in practice as soon as the coordinate has received any non-trivial gradient signal, but large enough to prevent division by zero at initialization. This is why the term is essential — without it, the first step would be infinite in any coordinate that received zero gradient, which is essentially every coordinate at initialization.
Adam: first moment + second moment + bias correction
Adam combines RMSProp's per-coordinate second-moment scaling with Polyak momentum's first-moment smoothing. The state is two moving averages:
The naive update would be , but and are initialized to zero, so at step they are biased toward zero — , not . Bias correction divides by and to recover an unbiased estimate of the moments:
The full Adam update with the standard defaults , , is:
Here is why the bias correction is necessary rather than cosmetic. At , , so . Dividing by gives back , an unbiased estimate. Without the correction, the first several updates would be systematically too small and the trajectory would lag behind the true gradient descent path. With , the bias shrinks below 10% by around ; with , the same happens for only around . Many practitioners find that Adam "doesn't work" on short runs because they never let the warm-up finish.
A geometric interpretation: Adam's update is approximately when is roughly constant across coordinates, because is the first-moment direction with each coordinate normalised to unit scale. The sign operation is what makes Adam robust to gradient magnitudes — a coordinate with a tiny but consistent gradient gets the same effective step as a coordinate with a huge but consistent gradient. SGD and SGD-with-momentum do not have this property: they step in proportion to the gradient, so large-gradient coordinates dominate the update.
The two hyperparameters and are not symmetric. controls the effective horizon of the first moment (the direction), and controls the effective horizon of the second moment (the per-coordinate scale). The first-moment horizon is typically much shorter than the second-moment horizon because the direction of the gradient changes quickly during training, while the per-coordinate scale is more stable. Setting (equal horizons) would make Adam behave more like SGD-with-momentum and less like RMSProp; setting and (the defaults) gives a much longer second-moment horizon than first-moment horizon, which is the regime in which Adam's per-coordinate rescaling dominates.
import numpy as np
def adam_step(params, grads, m, v, t, lr=1e-3, b1=0.9, b2=0.999, eps=1e-8):
"""One Adam step applied to a list of parameter arrays.
The bias-correction factor 1 - beta^t only makes sense per-parameter
when each parameter sees its own gradient at step t; here we use a
single global step counter so the warm-up is consistent.
"""
for i, (p, g) in enumerate(zip(params, grads)):
m[i] = b1 * m[i] + (1 - b1) * g # first moment
v[i] = b2 * v[i] + (1 - b2) * g * g # second moment
m_hat = m[i] / (1 - b1 ** t) # bias-corrected first moment
v_hat = v[i] / (1 - b2 ** t) # bias-corrected second moment
# The epsilon lives OUTSIDE the sqrt so the step is finite even when
# v_hat is exactly zero at initialization.
p -= lr * m_hat / (np.sqrt(v_hat) + eps)
return params, m, v
The default , is not magical; it reflects the fact that the first moment wants a shorter memory (so that recent gradients dominate the direction) while the second moment wants a much longer one (so that the per-coordinate scale estimate is stable). Mixing these two time scales is the point of Adam, and it is also the reason Adam sometimes fails where SGD-with-momentum succeeds: the second moment can be dominated by early, large gradients that no longer reflect the current loss surface.
AdamW: decoupling weight decay from the gradient
This is the section that distinguishes the 201-level treatment. Most deep learning libraries, when you "add weight decay" to Adam, do something like
This is coupled weight decay — the L2 penalty appears inside the gradient, which is then divided by . The result is that weights with large (consistently large gradients) get a much smaller effective weight-decay pull than weights with small . That is the opposite of what you want: you want weights that the model would like to grow large to be the ones that get pulled back toward zero.
Loshchilov & Hutter (2019) showed that the right way to combine weight decay with Adam is to subtract the decay term directly from the weights, after the Adam step:
The last term is the decoupled weight decay; it is independent of , so every weight gets the same proportional pull toward zero. This is the AdamW update. Let us unpack the algebraic step that makes the difference. In the coupled version, the L2 contribution to the gradient at iteration is . After the Adam scaling, the effective decay applied to coordinate is
In the decoupled version it is
The ratio is , which is huge for coordinates with small gradient magnitudes (think weights in early layers of a transformer, where the gradient signal is weak) and small for coordinates with large gradient magnitudes (think the classifier head, where gradients are large). AdamW fixes exactly this inversion of intent. Empirically, AdamW trains transformers and many other deep architectures more stably than Adam-with-coupled-decay, and it is the default in essentially every modern recipe.
The intuition for the difference is that Adam's per-coordinate rescaling is designed to make the optimization effective, not to make the regularization effective. When you put L2 weight decay inside the gradient, Adam's rescaling distorts the regularization in ways that are not the intent of the L2 prior. AdamW treats the optimization and the regularization as separate operations on the parameters: first Adam does its per-coordinate scaling to compute the right step, then weight decay shrinks the parameters proportionally.
def adamw_step(params, grads, m, v, t, lr=1e-3, b1=0.9, b2=0.999,
eps=1e-8, weight_decay=0.01):
"""AdamW: Adam with decoupled (L2) weight decay.
Notice the difference from adam_step: there is no L2 term inside `grads`.
The decay is subtracted from the parameters AFTER the Adam step, so it
is independent of the per-coordinate second moment.
"""
for i, (p, g) in enumerate(zip(params, grads)):
m[i] = b1 * m[i] + (1 - b1) * g
v[i] = b2 * v[i] + (1 - b2) * g * g
m_hat = m[i] / (1 - b1 ** t)
v_hat = v[i] / (1 - b2 ** t)
# Adam update, scaled by the current parameter itself so that the
# decay is proportional to |w|. Subtract the Adam step first.
p -= lr * m_hat / (np.sqrt(v_hat) + eps)
# Decoupled weight decay: a fixed proportional pull toward zero,
# not scaled by the second moment.
p -= lr * weight_decay * p
return params, m, v
A subtle point about the implementation: the weight decay is applied every step, not just at convergence. This means a parameter that has just been pushed away from zero by a large gradient will be pulled back toward zero on the next step by a fixed proportion, regardless of how large its gradient was. The decay rate is , so with and , every parameter shrinks by 1% per step. Over steps, a parameter shrinks by approximately from its unregularised trajectory. This is the implicit coupling between learning rate and weight decay that AdamW respects but coupled-decay Adam does not.
SGD vs Adam: when the defaults betray you
The default recommendation in most modern deep learning code is AdamW with , , , weight_decay = 0.01. This default was tuned on transformers and works well for that model class. It is not a universal default — for convolutional networks with batch normalisation, SGD with momentum and a larger initial learning rate () often reaches a better final loss.
Why does the default differ? The answer is in how each optimizer handles the gradient signal. In a transformer, the gradient magnitudes vary wildly across layers and across training steps; early-layer gradients are small, late-layer gradients are large, and the ratio can exceed 100. Adam's per-coordinate rescaling absorbs this variance and lets every parameter learn at roughly the same rate. SGD-with-momentum, by contrast, applies the same step size to every parameter, so the parameters with small gradients learn slowly while the parameters with large gradients learn fast.
In a convolutional network with batch normalisation, the per-layer gradient magnitudes are more uniform, because batch normalisation scales activations to have unit variance and the gradient signal becomes more homogeneous. Adam's rescaling is therefore less helpful, and the implicit regularization of SGD-with-momentum (tendency to find flat minima that generalize well) can dominate. Empirically, ResNets trained with SGD-with-momentum reach test accuracies 0.5-1.5 percentage points higher than the same architecture trained with Adam.
A second consideration is the batch size. Adam is more robust to small batch sizes because the per-coordinate second moment absorbs the noise; SGD-with-momentum on small batches produces noisy trajectories that require careful learning-rate scaling. If you are constrained to small batches (large models, long sequences), Adam is the safer default. If you can afford large batches (vision models on modern GPUs), SGD-with-momentum is worth considering for the final accuracy.
A third consideration is the training horizon. Adam converges faster in wall-clock steps but plateaus at a higher loss than SGD-with-momentum for very long training runs on some problems. The standard explanation is that Adam's adaptive rescaling keeps the iterate in sharper basins, while SGD-with-momentum's noise lets it find flatter basins. Whether this matters depends on whether you care about generalization (flatter is better) or about reaching a target loss quickly (faster is better).
The choice of optimizer also interacts with the batch size. When you increase the batch size by a factor of , the per-step gradient noise shrinks by , so the optimal learning rate scales roughly linearly with . For SGD this is straightforward; for Adam, the second-moment estimate absorbs some of the noise, so the optimal is less sensitive to batch size. This is why large-batch training of transformers (thousands of examples per step) is more tractable with Adam than with SGD.
A fourth consideration is gradient clipping. Transformers in particular produce gradient spikes during training, and the standard fix is to clip the gradient norm to some maximum value (often 1.0) before applying the optimizer. Gradient clipping interacts with Adam: the clip bounds the maximum step magnitude along any direction, while Adam's per-coordinate rescaling bounds the per-coordinate step. Both are useful, and the standard recipe applies both.
| Optimizer | When it shines | When it disappoints |
|---|---|---|
| SGD + momentum | Convolutional nets, large batches, long horizon | Transformers, small batches, ill-conditioned loss |
| Adam | RNNs, small batches, ill-conditioned loss | Long-horizon image classification, when SGD's flat-minima bias matters |
| AdamW | Transformers, fine-tuning, any task where weight decay helps | Same as Adam; the W in AdamW is what differentiates it from coupled-decay Adam |
| LARS / LAMB | Very large batches (>= 8k examples per step) | Small batches where per-coordinate scaling is overkill |
def lamb_step(params, grads, m, v, t, lr, b1, b2, eps, weight_decay):
"""LAMB: layer-wise adaptive moments for batch training.
Unlike Adam which rescales per-coordinate, LAMB rescales per-layer
by the ratio ||w|| / ||m_hat/sqrt(v_hat)||. This makes the optimizer
behave consistently across layers with different parameter scales
(think embedding layers vs classifier heads in transformers) and
allows very large batch training without divergence.
"""
for i, (p, g) in enumerate(zip(params, grads)):
m[i] = b1 * m[i] + (1 - b1) * g
v[i] = b2 * v[i] + (1 - b2) * g * g
m_hat = m[i] / (1 - b1 ** t)
v_hat = v[i] / (1 - b2 ** t)
update = m_hat / (np.sqrt(v_hat) + eps)
# Trust ratio: ratio of parameter norm to update norm, bounded.
trust = np.linalg.norm(p) / (np.linalg.norm(update) + 1e-8)
p -= lr * trust * update - lr * weight_decay * p
return params, m, v
Learning-rate schedules
A fixed learning rate is rarely optimal. The intuition is that you want a large step early — when the iterates are far from any minimum — and a small step late — when the iterates are close and a large step would oscillate. There are several ways to encode this, and the right choice depends on the optimiser and the training horizon.
Step decay drops by a factor (typically 0.1) at fixed epochs. It is simple and works well when you know the right schedule in advance. For ResNets, dropping the learning rate at 1/3 and 2/3 of the training horizon by a factor of 10 is a standard recipe.
Cosine annealing decreases smoothly according to
where is the total step budget. It is the standard default in modern vision recipes. The smooth decrease avoids the abrupt behaviour of step decay and tends to give slightly better final accuracy.
Warmup-then-decay linearly increases from 0 to the peak over the first few hundred steps, then decays. Warmup matters for Adam and AdamW because the bias-correction denominator is far from 1 at small — without warmup, the early effective step is much smaller than intended, and the iterates take a few epochs to start moving at the right speed. For transformers, warmup is essentially mandatory.
A subtle point: schedules interact with adaptive optimizers in a way that does not happen with SGD. Because Adam normalizes the step by , multiplying by a constant is not equivalent to multiplying the gradient by ; the per-coordinate scaling absorbs part of the change. This is why warmup matters even with Adam: the early steps need a smaller effective step not because the raw gradient is large, but because the second-moment estimate has not yet converged.
A second subtle point: the interaction between schedule and batch size. When you increase the batch size by a factor of , the optimal initial learning rate scales roughly linearly with for SGD but sublinearly for Adam. This means a schedule tuned for batch size 32 with SGD cannot be directly transferred to batch size 256 with Adam — both the peak learning rate and the warmup duration need to be re-tuned. The standard remedy is to use a learning-rate finder (a few epochs of exponentially increasing , then take the largest at which the loss still decreases).
Step decay drops by a factor (typically 0.1) at fixed epochs. It is simple and works well when you know the right schedule in advance.
Cosine annealing decreases smoothly according to
where is the total step budget. It is the standard default in modern vision recipes.
Warmup-then-decay linearly increases from 0 to the peak over the first few hundred steps, then decays. Warmup matters for Adam and AdamW because the bias-correction denominator is far from 1 at small — without warmup, the early effective step is much smaller than intended, and the iterates take a few epochs to start moving at the right speed.
A subtle point: schedules interact with adaptive optimizers in a way that does not happen with SGD. Because Adam normalizes the step by , multiplying by a constant is not equivalent to multiplying the gradient by ; the per-coordinate scaling absorbs part of the change. This is why warmup matters even with Adam: the early steps need a smaller effective step not because the raw gradient is large, but because the second-moment estimate has not yet converged.
Diagnosing optimiser pathologies
When a training run goes wrong, the loss curve usually tells you which knob is mis-sized, if you know how to read it. Below is a worked example of interpreting a typical pathological curve.
A loss curve that drops for the first 100 steps and then oscillates with a period of a few hundred steps, with the amplitude of the oscillation growing, indicates an effective step size that is on the edge of stability. The model is close to a region of the loss surface where the curvature along the update direction is smaller than the step size, so the iterates overshoot and then have to be pulled back, then overshoot again, then pull back. The remedies, in order of how often they work, are: (a) halve the learning rate; (b) add gradient clipping at some maximum norm (often 1.0 for transformers); (c) switch from SGD to Adam with a much smaller initial learning rate ( instead of ).
A loss curve that is flat for thousands of steps and then suddenly drops indicates that the optimizer was stuck in a saddle point or a flat basin and finally escaped. This is not a pathology — it is the expected behaviour of SGD on a non-convex loss surface. The remedy is patience, not a hyperparameter change. A common mistake is to lower the learning rate in response, which slows the escape; the right move is to wait, or to switch to Adam, whose second-moment estimate makes the effective step larger in flat regions.
A loss curve that diverges immediately (loss goes to NaN or infinity within the first 50 steps) almost always means an exploding gradient. Check whether any layer's output has a non-finite value, or whether any parameter is being multiplied by a number larger than 1.0 at every step. The remedies are: gradient clipping, a smaller initialisation variance, layer normalisation, or a learning rate that is one or two orders of magnitude smaller.
A loss curve that decreases for a while, then increases slowly, suggests overfitting on the training set. The optimizer is correctly fitting the training data and the model is now memorizing the labels. The remedies are: stronger regularisation (higher weight decay), dropout, data augmentation, or a smaller model. The optimizer is doing its job; the model class is too expressive for the dataset.
A loss curve that decreases monotonically but plateaus at a high value (e.g. 0.5 cross-entropy for a binary classification problem) suggests the model is capacity-limited. The optimizer has found the best parameters it can find, but the model cannot fit the training data any better. The remedies are: a larger model, better features, or a longer training schedule. A common mistake is to lower the learning rate in response, which makes the plateau more visible but does not address the underlying capacity limit.
def train_step_with_clipping(params, grads, m, v, t, lr, b1, b2, eps,
weight_decay, max_norm=1.0):
"""AdamW step with per-parameter gradient clipping.
The clipping is applied to the gradient magnitudes, not to the parameter
update: we compute the total norm of the gradient vector, and if it
exceeds max_norm we scale every gradient by max_norm / total_norm.
This preserves the direction of the update while bounding the step.
"""
total_norm = 0.0
for g in grads:
total_norm += float(np.sum(g * g))
total_norm = np.sqrt(total_norm)
clip_coef = max_norm / (total_norm + 1e-6)
if clip_coef < 1.0:
grads = [g * clip_coef for g in grads]
return adamw_step(params, grads, m, v, t, lr, b1, b2, eps, weight_decay)
The interaction between optimiser and initialisation is also worth understanding. Xavier initialisation keeps the variance of activations roughly constant across layers, and Kaiming initialisation does the same for ReLU networks. Without proper initialisation, the gradients at the early layers can be vanishingly small or explosively large, and the optimiser spends the first few epochs correcting for the initialisation rather than learning the task. Modern practice — layer normalisation in transformers, batch normalisation in convolutional networks — largely removes the need for careful initialisation by renormalising activations at every layer.
A related diagnostic is the gradient norm per layer. After a few hundred steps of a healthy training run, the per-layer gradient norms should be roughly equal (or at least no layer should be many orders of magnitude larger or smaller than the others). If one layer has gradient norm while the others are around 1, that layer is either misinitialised or has an exploding activation. Gradient clipping bounds the global gradient norm but does not address per-layer imbalance — for that, you need per-layer normalisation.
Practical recipe
The defaults below are what you should reach for first; tune only when you have a measured reason.
| Optimizer | Defaults | Use when |
|---|---|---|
| SGD + momentum | , | ResNets on image classification; small models where you want explicit control |
| Adam | , , | Quick experiments, RNNs, anything where SGD stalls |
| AdamW | to , , , weight_decay = 0.01 | Transformers, modern deep nets; the current safe default |
A divergence in the first 200 steps almost always means is too large or the gradient is exploding (look for NaN). A flat loss that does not decrease past a certain point almost always means is too small or the model is capacity-limited (look at the gap to a larger model). An oscillating loss curve almost always means is too large for the current batch size, or that the data ordering is causing correlated batches — shuffle more aggressively or increase the batch.
The training loss curve is informative, but the validation loss curve is the one that matters for model selection. A gap between training and validation loss that grows over time indicates overfitting; the remedy is stronger regularisation (higher weight decay, more dropout, more data augmentation). A gap that stays small but both losses are high indicates underfitting; the remedy is a larger model, better features, or a longer training schedule. The intersection of these two regimes — where training loss continues to fall but validation loss plateaus or rises — is where early stopping helps.
The optimizer is rarely the bottleneck of a real ML project. Once you have picked AdamW with sensible defaults and a cosine schedule, the next 80% of your model's performance will come from data quality, feature engineering, and capacity — topics in later lessons.
Why the optimiser is rarely the bottleneck
The most important property of an optimizer is predictability. Given a fixed compute budget and a fixed model architecture, a good optimizer reaches a known loss in a known number of steps, with a known memory footprint, and the result is reproducible across runs. Adam and AdamW have this property; so does SGD-with-momentum when paired with a cosine schedule. Exotic optimizers (LARS, Shampoo, SOAP, Lion) have higher peak performance on some tasks but lower predictability — they require more careful tuning and their behaviour on out-of-distribution problems is less well understood.
In a typical ML project, the optimizer contributes a few percentage points of accuracy at most. The much larger contributions come from: (1) data quality and quantity, (2) feature engineering and pipeline hygiene (see lesson 07), (3) model architecture and capacity, (4) training horizon and learning-rate schedule. If your model is underperforming, the optimizer is almost never the first place to look.
The optimizer also interacts with the rest of the system in subtle ways. Changing the optimizer changes the implicit regularization; changing the batch size changes the noise scale; changing the schedule changes the effective regularization strength. A "fair comparison" between optimizers requires holding all of these fixed, which is harder than it sounds because each optimizer has its own preferred settings. The literature on optimizer comparisons is full of confounding factors.
For most practitioners, the right approach is: pick AdamW with the standard defaults, use a cosine schedule with a small warmup, train for a fixed budget of epoch-equivalents, and stop when the validation loss stops improving. Spend the rest of your time on data and features. If you have a measured reason to try a different optimizer — say, SGD-with-momentum is consistently beating AdamW on your convolutional tasks — switch and document the comparison. Otherwise the default is fine.
The bias-variance tradeoff also enters the optimizer choice. Adam's per-coordinate rescaling produces sharper minima (in the curvature sense) than SGD-with-momentum on some problems, which can hurt generalization. The standard explanation is that Adam's adaptive learning rate makes the iterates aggressive about following the gradient in any direction with consistent signal, while SGD's noise forces the iterates to commit to basins that are robust to small perturbations in the data. This is not a universal rule — some tasks prefer sharper minima — but it is the default assumption in the deep learning literature. If your validation loss is consistently higher than your training loss by a wide margin, switching from Adam to SGD-with-momentum is one of the first things to try.
Batch size, gradient noise, and the implicit temperature
The learning rate is not the only knob that sets how far the iterates travel per step. The mini-batch size does too, and the mechanism is worth deriving because it explains why "just use a bigger batch" is not free.
The stochastic gradient is an unbiased estimator of the full gradient, so in expectation the iterate performs ordinary gradient descent. But the fluctuation around that expectation is what matters. For squared loss with per-example gradients , the gradient noise scale of a batch of size is approximately
where is the variance of the per-example gradient. The Stationary Distribution Approximation treats SGD as sampling from a distribution around the minimum whose covariance is , with the Hessian at the minimum. Read that as: the noise scale is set by the ratio .
Two consequences follow, and both are routinely observed. First, doubling the batch size and doubling the learning rate holds fixed and so roughly preserves the noise scale — which is why large-batch training uses a scaled learning rate plus a warmup, rather than silently shrinking the effective temperature. Second, a constant learning rate does not settle: the iterates keep bouncing inside a ball of radius around the minimum forever. Only a decaying schedule — cosine or linear to zero — actually anneals to a point. A "converged" run at constant is a run whose parameters are still being randomized within a ball whose size you chose.
Linear scaling has a limit, though. Beyond a certain batch size the approximation breaks because the gradient's heavy tail — a few examples with enormous per-example gradients — dominates the variance, and the variance stops falling like . That is why SGD is paired with a momentum-like term that is applied outside the batch sum: a large batch is computationally parallel but statistically noisy, and the fix for statistical noise is not more parallelism.
def linear_warmup_cosine(step, total_steps, warmup_steps, base_lr, min_lr=0.0):
"""The standard transformer schedule, and the reason it exists.
Warmup is not a convenience: at step 0 the second-moment estimate v_t is near zero,
so m_t / sqrt(v_t) is a ratio of two near-zero numbers and the update's magnitude is
essentially arbitrary. Ramp the learning rate in linearly so those early steps stay
bounded, then anneal with a cosine so the stationary distribution's radius shrinks
to zero.
"""
if step < warmup_steps:
return base_lr * (step + 1) / max(1, warmup_steps)
progress = (step - warmup_steps) / max(1, total_steps - warmup_steps)
progress = min(1.0, progress)
return min_lr + 0.5 * (base_lr - min_lr) * (1.0 + math.cos(math.pi * progress))
The preconditioner view, and why Adam is diagonal
Everything in this lesson can be stated in one frame: gradient descent picks a direction and a step size, but the loss surface is rarely isotropic, so the gradient direction is usually not the direction of steepest descent in the geometry that matters. The fix is a preconditioner , giving the update .
The ideal is the inverse Hessian , because Newton's method — which is exactly that — converges quadratically. The catch is that is for parameters: for a 7-billion-parameter transformer, forming it or inverting it is out of the question, and even a low-rank or block-diagonal approximation is usually too expensive to update every step.
Adam is the cheap approximation. Its update is, to first order, restricted to the class of diagonal preconditioners: each coordinate is divided by its own recent gradient scale and nothing else. AdaGrad is the same idea with a running sum of squares and no exponential average; RMSProp is the same idea with an exponential average and no momentum; Adam is both. This is the cleanest way to see the family — they are not different algorithms, they differ only in the estimator used for the second moment and whether a first moment is carried.
What a diagonal preconditioner cannot do is rotate coordinates. If two loss directions are strongly correlated — one common case is a redundant pair of input features, where the gradient consistently pushes along the line — a diagonal shrinks both by their own scale but cannot eliminate the coupling, so the iterate still zig-zags along the ill-conditioned direction. Full-matrix methods (Shampoo, SOAP, K-FAC) attack exactly this by estimating blocks of the curvature in a transformed basis, and the reason they are not the default is cost, not principle.
Reading a training run: four numbers that diagnose most failures
Before changing anything about the optimizer, read these off the run. They localise the problem faster than any hyperparameter sweep.
Gradient norm at init. If is enormous ( for a cross-entropy model), the loss surface near initialisation is almost vertical and no learning rate will rescue it. The cause is almost always unscaled inputs: a feature with range makes the logits enormous, the softmax saturates, and the gradient carries a factor of that scale. Fix the data, not the optimizer.
Train loss after one epoch. If it is still at initialisation, the model is not learning at all — a dead ReLU, a vanishing gradient, a bug in the loss, or three orders of magnitude too small. This is a correctness bug masquerading as a tuning problem, and no schedule will find it.
The train/val gap at the end. A large gap is overfitting, and the remedy is regularization (lesson 06), not the optimizer. A small gap with high error everywhere is underfitting, and the remedy is capacity or features. Reading the gap correctly is what stops people from tuning the optimizer when the answer is in lesson 06 or 07.
Loss-curve noise amplitude. If training loss oscillates by more than a few percent between logged steps, the noise scale is too large. Either decay the learning rate faster or increase the batch. If instead the curve is smooth but the validation curve is noisy step to step, that is the evaluation set being too small to measure the quantity you care about — a resolution problem, not a training problem.
Key Takeaways
- Plain SGD fails because it cannot handle ill-conditioned loss surfaces (no per-coordinate scaling) and cannot exploit the temporal structure of mini-batch gradients (no smoothing of the update direction). Every modern optimizer addresses at least one of these gaps.
- Momentum carries the iterate through shallow basins and damps oscillations across valley walls, but does not change the per-coordinate scale of the update.
- Adam combines first-moment smoothing (Polyak momentum) with second-moment scaling (RMSProp) and bias correction so the iterates have the right expected direction and magnitude even at small step counts.
- AdamW is not "Adam with an extra term" — it is the same Adam update with weight decay subtracted from the parameters directly, independent of the second-moment estimate. The decoupling is what makes the regularization behave intuitively under per-coordinate scaling.
- A learning-rate schedule is part of the optimizer, not a separate concern. Warmup matters for Adam because the bias correction takes time to converge.
- Diagnose loss-curve pathologies by reading the shape: divergence is too-large or exploding gradient; flat-then-stuck is too-small or capacity limit; oscillation is too-large or correlated batches.
- For most real projects, the optimizer is rarely the bottleneck — pick AdamW with the standard defaults, use a cosine schedule with warmup, and spend your time on data quality, feature engineering, and capacity.