Multiple Linear Regression
One model, many features. The normal equation w = (XᵀX)⁻¹Xᵀy solves it in closed form — a formula GATE wants you to apply to tiny matrices, not to derive.
What you'll learn
- Extending the line to many features: the model ŷ = Xw with a bias column
- The normal equation w = (XᵀX)⁻¹Xᵀy as a formula to apply, not derive
- Reading off coefficients, and the matrix shapes of X, w, and y
- Why XᵀX must be invertible (full column rank) for the closed form to exist
Before you start
Last lesson left the line straining against its own limit: a single slope, when a house price answers to size and location and age and room count, all at once. One number cannot carry four influences. So the line has to grow a weight for each feature — and, happily, there is still a single closed-form solution waiting on the other side.
Multiple linear regression is simple regression with the lone slope generalised to one weight per feature. The single line becomes a weighted sum of inputs, and the tidy scalar hand-formula Σxy / Σx² becomes one clean matrix equation. The bookkeeping moves from numbers to matrices, but the idea — minimise the squared error — does not budge an inch.
The model in matrix form
Stack your data so each row of the matrix X is one example and each column is one feature. Then add a leading column of 1s — the bias column — so the intercept rides along as just another weight. The predictions for every row at once become ŷ = Xw, a single matrix-vector product:
Each entry of w is a coefficient: holding the other features fixed, it is the change in the prediction per one-unit increase in that feature. The bias weight — the coefficient on the column of 1s — is the intercept the previous lesson called c.
The normal equation — a formula to apply
Minimising the squared error Σ(yᵢ − ŷᵢ)² over all the weights at once has a single closed-form answer, the normal equation:
GATE wants you to apply this, not derive it. The mechanical recipe is:
- Build the small
d × dmatrixXᵀX. - Build the vector
Xᵀy. - Invert the matrix.
- Multiply.
For a 2 × 2 matrix the inverse is the familiar (1/det) times the swap-and-negate pattern, so the whole thing is doable by hand.
Why not invert X directly?
Now notice what the formula pointedly does not say. The shortcut everyone reaches for first does not exist: you cannot write w = X⁻¹y. X is n × d. In any real fitting problem there are far more rows than columns. A rectangular matrix has no inverse at all, so there is simply nothing there to invert.
The XᵀX sandwich is precisely the repair. Multiplying by Xᵀ on the left folds the tall X down into a small square d × d matrix, and that one can be inverted.
(In the degenerate case where X does happen to be square and invertible, the formula quietly collapses back to w = X⁻¹y. This follows from (XᵀX)⁻¹Xᵀ = X⁻¹(Xᵀ)⁻¹Xᵀ = X⁻¹ — but that means you have exactly as many equations as unknowns, and nothing left to fit.)
How GATE asks this
Either a MCQ that asks you to recognise the normal equation w = (XᵀX)⁻¹Xᵀy (or pick the correct matrix shapes), or a NAT that hands you a tiny X and y and asks for one coefficient. With a 2 × 2 matrix XᵀX the inverse is one line of arithmetic. The graders keep the numbers clean on purpose, so the matrix algebra — not the calculator work — is what gets tested.
Worked example — recover a line from three points
Fit
ŷ = w₀ + w₁·xto the points(1, 3),(2, 5),(3, 7)using the normal equation. Findw₀andw₁.
With a bias column, the design matrix X — that is just the data laid out with one row per example and one column per feature — together with the weight vector and the target, are:
⎡ 1 1 ⎤ ⎡ w₀ ⎤ ⎡ 3 ⎤
X = ⎢ 1 2 ⎥ w = ⎣ w₁ ⎦ y = ⎢ 5 ⎥
⎣ 1 3 ⎦ ⎣ 7 ⎦
Build XᵀX (a 2×2) and Xᵀy (a 2-vector). Each entry of XᵀX is a dot product of two columns of X:
- The top-left is column-1 with itself (
1·1+1·1+1·1). - The off-diagonal is column-1 with column-2 (
1·1+1·2+1·3). - The bottom-right is column-2 with itself (
1·1+2·2+3·3):
⎡ 1+1+1 1+2+3 ⎤ ⎡ 3 6 ⎤
XᵀX = ⎣ 1+2+3 1+4+9 ⎦ = ⎣ 6 14 ⎦
⎡ 3 + 5 + 7 ⎤ ⎡ 15 ⎤
Xᵀy = ⎣ 1·3 + 2·5 + 3·7 ⎦ = ⎣ 34 ⎦
Invert the 2×2 (determinant = 3·14 − 6·6 = 42 − 36 = 6):
1 ⎡ 14 −6 ⎤
(XᵀX)⁻¹ = ─── ⎣ −6 3 ⎦
6
1 ⎡ 14·15 − 6·34 ⎤ 1 ⎡ 210 − 204 ⎤ 1 ⎡ 6 ⎤ ⎡ 1 ⎤
w = ─── ⎣ −6·15 + 3·34 ⎦ = ─── ⎣ −90 + 102 ⎦ = ─── ⎣ 12 ⎦ = ⎣ 2 ⎦
6 6 6
So w₀ = 1, w₁ = 2 — the model is ŷ = 1 + 2x. The same steps in NumPy confirm it:
import numpy as np
X = np.array([[1, 1],
[1, 2],
[1, 3]], dtype=float) # bias column + one feature
y = np.array([3, 5, 7], dtype=float)
XtX = X.T @ X
Xty = X.T @ y
w = np.linalg.inv(XtX) @ Xty
print("XtX =", XtX.tolist())
print("Xty =", Xty.tolist())
print("w =", w.tolist())
XtX = [[3.0, 6.0], [6.0, 14.0]]
Xty = [15.0, 34.0]
w = [1.0, 2.0]
Check it against the data: at x = 1, 2, 3 the model gives 3, 5, 7, matching every point exactly. These three points are perfectly collinear, so the residuals are all zero and the fit is exact — just as the prediction prompt suggested.
In one breath
Multiple linear regression stacks the data into a design matrix X (rows = samples, columns = features, plus a leading bias column for the intercept) and predicts all rows at once as ŷ = Xw.
Minimising the squared error over every weight together gives the single closed-form normal equation w = (XᵀX)⁻¹Xᵀy. You apply it by forming the small d×d matrix XᵀX, inverting it, and multiplying by Xᵀy. It only works when XᵀX is invertible, i.e. when no feature column is a linear combination of the others.
Practice
Quick check
A question to carry forward
The normal equation is a marvel: one shot, no iteration, the exact best weights. But read its cost again — it forms XᵀX and then inverts it. Inverting a d × d matrix costs on the order of d³ operations, and the inverse only exists when the columns are independent. With a handful of features that is nothing. With ten thousand features, or columns that very nearly collide, the one-shot formula becomes slow, memory-hungry, or simply undefined.
So we need a second route to the same minimum — one that never forms an inverse, never even builds XᵀX, but instead creeps toward the best weights a little at a time. Here is the thread onward: if you are standing somewhere on the bowl-shaped cost surface and want to reach its bottom, which direction is downhill, and how far should you step? What single update rule turns that into an algorithm?
Practice this in an interview
All questionsUnder full column rank, OLS sets the gradient of the squared-error objective to zero, giving the normal equations and the unique coefficient vector β = (XᵀX)⁻¹Xᵀy. In rank-deficient or numerical settings, use the pseudoinverse or a least-squares solver rather than explicitly forming the inverse.
Use gradient descent when the feature matrix is too wide, sparse, or continuously arriving for a direct least-squares solve to fit comfortably in memory, and use mini-batch or stochastic updates when you need online learning. For a modest, fixed, well-conditioned dense dataset, a direct solver is usually simpler and faster; there is no universal feature-count cutoff.
Scale features for distance-based models, PCA, neural networks, and regularized linear or logistic models because magnitudes affect geometry, optimization, or coefficient penalties. Tree-based models generally do not need scaling because threshold splits preserve feature order; ordinary unregularized linear models may work without it, but scaling often improves numerical conditioning.
The usual five assumptions are linearity, independent errors, constant error variance, normal residuals for exact small-sample inference, and no perfect multicollinearity; a technically complete answer also requires zero conditional mean. Each violation harms a different part of OLS: fit and bias, efficiency, standard errors, inference, or coefficient identifiability.