Ridge Regression & Regularization
Ordinary least squares can overfit. Ridge adds an L2 penalty that shrinks the weights, trading a little bias for a lot less variance.
What you'll learn
- Why unconstrained least squares overfits, and how a penalty term fixes it
- Ridge = L2 penalty: minimize Σ(yᵢ − ŷᵢ)² + λ·‖w‖²
- Ridge is L2, lasso is L1 — the single distinction GATE checks
- Larger λ shrinks weights toward zero: more bias, less variance
- Computing a regularized loss to a clean number
Before you start
Last lesson closed on an uncomfortable truth: our optimizers are too good. Both the normal equation and gradient descent drive the training loss to its floor — and at that floor, with many features or noisy data, sit huge, jittery weights that fit the noise and swing wildly on anything new. That is overfitting, and you cannot optimize your way out of it, because the trap is the very bottom of the bowl you are descending.
So we change the bowl. Instead of asking the model only to fit the data, we add a second clause to its objective: and keep your weights small. The model now has to balance two pulls — match the data, yet stay modest — and the modesty clause tugs every weight toward zero, smoothing the fit. That added clause is regularization, and its most common form is ridge regression.
The L2 penalty
Ridge minimises the usual squared error plus a penalty on the size of the weight vector, scaled by a knob λ (lambda):
- The first term,
Σ(yᵢ − ŷᵢ)², is ordinary least squares — make the predictions match. ‖w‖² = Σ wⱼ²is the L2 penalty: the sum of the squares of the weights.λ ≥ 0controls the trade-off. Atλ = 0you recover plain least squares; asλgrows, the penalty dominates and every weight is pulled toward zero (though, for ridge, never exactly to zero).
One naming point GATE leans on hard: ridge uses the L2 norm; lasso uses the L1 norm (Σ|wⱼ|, the sum of absolute values). Lasso can drive weights to exactly zero, performing feature selection; ridge only shrinks them. Do not swap the two.
Slide λ up and watch every coefficient shrink. Toggle between L2 (ridge — a smooth shrink, nothing reaching zero) and L1 (lasso — small weights snapping to zero):
L1 zeroes out features. L2 just shrinks.
How GATE asks this
Two recurring shapes. As an MCQ, it asks the direction of the effect: what happens to bias and variance as you raise λ. GATE DA 2026 posed exactly this (Q37), and the correct statement was that the regularizer increases bias and decreases variance. As a NAT, it gives you weights, a data-fit error, and λ, and asks for the total regularized objective — pure substitution (GATE DA 2026 Q55 did this with an MAE data term).
Worked example
Part A — the direction (GATE DA 2026). Two words first, in plain terms: bias is how far a model’s predictions sit from the truth on average, and variance is how much those predictions jump around when you retrain on a different sample of data.
As the regularization coefficient λ rises, the weights are squeezed toward zero, so the model becomes simpler and stiffer. A stiffer model varies less from one training set to the next (variance decreases), but systematically misses more structure (bias increases). Push λ far enough and the model is so stiff it underfits — every weight near zero, predicting nearly the same value everywhere. The 2026 answer was exactly that pair: bias ↑, variance ↓.
One word is doing double duty here, and it trips almost everyone: the “bias” that goes up with λ is not the bias term b — the intercept in ŷ = w·x + b. They are unrelated ideas wearing the same name; one is a component of prediction error, the other is a single learnable parameter. In fact the intercept is normally left out of the penalty altogether: ‖w‖² sums over the feature weights only, since shrinking b would just drag every prediction toward zero for no good reason.
Part B — a regularized loss (clean numbers). Take the ridge objective J = (squared error on the data) + λ · ‖w‖² with weights w = (3, −1), a data-fit squared error of 4, and λ = 0.5:
‖w‖² = 3² + (−1)² = 9 + 1 = 10
penalty = λ · ‖w‖² = 0.5 · 10 = 5
J = data error + penalty = 4 + 5 = 9
The same steps in Python:
w = [3, -1]
data_error = 4
lam = 0.5
l2_norm_sq = sum(wi**2 for wi in w) # 9 + 1 = 10
penalty = lam * l2_norm_sq # 0.5 * 10 = 5
J = data_error + penalty # 4 + 5 = 9
print(f"||w||^2 = {l2_norm_sq}")
print(f"penalty = {penalty}")
print(f"J = {J}")
||w||^2 = 10
penalty = 5.0
J = 9.0
So the total regularized objective is J = 9. Notice the penalty (5) is larger than the data error (4): with λ = 0.5 and weights this size, the model really is under pressure to shrink them.
In one breath
Ordinary least squares can overfit by driving the weights huge to chase noise, so ridge regression adds an L2 penalty to the objective — minimise Σ(yᵢ − ŷᵢ)² + λ·‖w‖², where ‖w‖² = Σwⱼ² and λ ≥ 0 sets how hard the weights are pulled toward zero (λ = 0 is plain OLS); raising λ shrinks the weights, lowering variance at the cost of higher bias, and the one naming trap is that ridge is L2 while lasso is L1 (Σ|wⱼ|), the version that can zero weights out entirely.
Practice
Quick check
A question to carry forward
Ridge gave us one precise dial: turn λ up, and variance falls while bias rises. We have been saying those two words as if their meaning were settled — but they are the deepest pair in all of modelling, and ridge only nudges one instance of a tug that every model feels.
Why can you never just drive both to zero? A model simple enough to be steady across datasets (low variance) is too rigid to capture the truth (high bias); a model flexible enough to capture every wrinkle (low bias) reshapes itself wildly from one sample to the next (high variance). Here is the thread onward: what exactly are bias and variance, how do they split the total error of any model, and where does their sum bottom out — the sweet spot every technique in this chapter is secretly chasing?
Practice this in an interview
All questionsL1 adds the sum of absolute coefficient values to the loss, which drives some coefficients to exactly zero and performs implicit feature selection. L2 adds the sum of squared coefficients, which shrinks all weights proportionally but rarely zeroes any out. Lasso is preferred when you suspect only a few features matter; Ridge is preferred when most features contribute small effects.
Ridge regression is maximum a posteriori estimation for a linear model with Gaussian observation noise and a zero-mean Gaussian prior on the coefficients. The regularization strength is the noise variance divided by the prior variance, subject to the scaling convention used in the Ridge objective.
L1 and L2 both reduce variance by penalizing large coefficients, usually adding some bias in return. L1 can set coefficients exactly to zero, while L2 shrinks correlated features more smoothly; choose based on whether sparse selection or stable prediction matters, and consider elastic net when both matter.
L2 regularisation adds a penalty equal to the sum of squared weights multiplied by a coefficient λ to the loss function, which encourages the optimiser to keep weights small. This penalises large, specialised weights and pushes the model toward simpler solutions that generalise better. In SGD it is equivalent to shrinking weights by a constant factor each step (hence weight decay), though in Adam the two diverge — requiring AdamW for correct decoupled decay.