Gauss–Markov Theorem: Why OLS Is BLUE
A careful proof of the Gauss–Markov theorem: the assumptions OLS needs, what BLUE really means, and why normal errors are not required.
Ordinary least squares gives us a line. The Gauss–Markov theorem tells us exactly what is special about the coefficients of that line—and, just as importantly, what the theorem does not promise.
The short version is famous:
Under the Gauss–Markov assumptions, the ordinary least-squares estimator is BLUE: the Best Linear Unbiased Estimator.
That sentence is compact enough to memorize and vague enough to misuse. “Best” does not mean best for every purpose. “Linear” describes the estimator’s dependence on the observed responses, not necessarily a straight-line relationship between every variable. “Unbiased” is a repeated-sampling property, not a claim that one fitted line is correct. And normal errors are not part of the theorem.
Let’s build the result carefully, first for simple regression and then in matrix form so that the proof covers the intercept, the slope, and multiple regression all at once.
The model and the assumptions
For simple regression with an intercept, write
$$y_i = \beta_0 + \beta_1 x_i + \varepsilon_i, \qquad i=1,\ldots,n.$$The fixed-design version treats the observed $x_i$ values as fixed. A more general random-design version conditions on the entire design matrix $X$. Conditional on $X$, the Gauss–Markov assumptions are:
- Linear in the unknown coefficients: $y=X\beta+\varepsilon$.
- Full column rank: the columns of $X$ are linearly independent. In simple regression with an intercept, this means that not all $x_i$ are equal.
- Zero conditional mean: $E(\varepsilon\mid X)=0$.
- Spherical conditional covariance:
The last line combines two conditions: every error has the same conditional variance, and distinct errors are conditionally uncorrelated.
This formulation also fixes three common misunderstandings.
First, $E(\varepsilon_i\mid X)=0$ does not say that the residual cloud must look symmetric. A distribution can have mean zero and still be skewed. The condition says that, after conditioning on the regressors, the systematic part of the response has already been captured by $X\beta$.
Second,
$$\operatorname{Cov}(\varepsilon_i,\varepsilon_j\mid X)=0$$does not generally imply that $\varepsilon_i$ and $\varepsilon_j$ are independent. Independence implies zero covariance when the relevant moments exist, but the converse needs extra structure—for example, joint normality.
Third, a “linear model” is linear in $\beta$. The columns of $X$ may contain $x^2$, interactions, or other known transformations. What matters for Gauss–Markov is the form $X\beta$, not whether the fitted curve is visually straight in the original variable.
What OLS computes
OLS chooses $\hat\beta$ to minimize the residual sum of squares
$$\lVert y-Xb\rVert^2.$$Because $X$ has full column rank, $X^\mathsf{T}X$ is invertible, and the unique solution is
$$\hat\beta=(X^\mathsf{T}X)^{-1}X^\mathsf{T}y.$$For simple regression, define
$$S_{xx}=\sum_{i=1}^n(x_i-\bar x)^2.$$Then the familiar coefficients are
$$\hat\beta_1=\frac{\sum_{i=1}^n(x_i-\bar x)(y_i-\bar y)}{S_{xx}}, \qquad \hat\beta_0=\bar y-\hat\beta_1\bar x.$$An estimator is the random rule that maps a sample to a value. An estimate is the numerical value that rule produces for one observed sample. Thus $\hat\beta$ is an estimator before the data are observed and an estimate after the numbers are plugged in.
Why OLS is unbiased
Let
$$M_0=(X^\mathsf{T}X)^{-1}X^\mathsf{T},$$so that $\hat\beta=M_0y$. Conditional on $X$,
$$\begin{aligned} E(\hat\beta\mid X) &=M_0E(y\mid X)\\ &=M_0X\beta\\ &=(X^\mathsf{T}X)^{-1}X^\mathsf{T}X\beta\\ &=\beta. \end{aligned}$$Every component of $\hat\beta$ is therefore unbiased. That includes both $\hat\beta_0$ and $\hat\beta_1$ in the simple model.
Its conditional covariance matrix is
$$\begin{aligned} \operatorname{Cov}(\hat\beta\mid X) &=M_0\operatorname{Cov}(y\mid X)M_0^\mathsf{T}\\ &=\sigma^2M_0M_0^\mathsf{T}\\ &=\sigma^2(X^\mathsf{T}X)^{-1}. \end{aligned}$$For simple regression, this gives
$$\operatorname{Var}(\hat\beta_1\mid X)=\frac{\sigma^2}{S_{xx}},$$$$\operatorname{Var}(\hat\beta_0\mid X)=\sigma^2\left(\frac1n+\frac{\bar x^2}{S_{xx}}\right),$$and
$$\operatorname{Cov}(\hat\beta_0,\hat\beta_1\mid X)=-\frac{\sigma^2\bar x}{S_{xx}}.$$So the intercept has not disappeared. It is already part of the same vector calculation.
What “best” means
We need to be precise about the competition.
Consider any other estimator of the full coefficient vector that is linear in $y$:
$$\tilde\beta=Ay,$$where $A$ may depend on the fixed—or conditioned-on—design $X$, but not on $y$.
For $\tilde\beta$ to be unbiased for every possible $\beta$,
$$E(\tilde\beta\mid X)=AX\beta=\beta$$must hold for every $\beta$. Therefore
$$AX=I_p.$$The phrase “best linear unbiased” means that, among all matrices $A$ satisfying this condition, OLS has the smallest covariance matrix in the positive-semidefinite ordering. Equivalently, every linear combination $c^\mathsf{T}\hat\beta$ has variance no greater than the corresponding $c^\mathsf{T}\tilde\beta$.
That is narrower—and more useful—than saying “OLS is the best estimator in existence.” Biased estimators can have lower mean-squared error in some settings, and nonlinear estimators are outside this comparison.
The complete BLUE proof
Set
$$D=A-M_0.$$Both estimators are unbiased, so
$$DX=(A-M_0)X=I_p-I_p=0.$$This identity kills the cross terms that are about to appear. In particular,
$$M_0D^\mathsf{T}=(X^\mathsf{T}X)^{-1}X^\mathsf{T}D^\mathsf{T}=(X^\mathsf{T}X)^{-1}(DX)^\mathsf{T}=0,$$and likewise $DM_0^\mathsf{T}=0$.
Now compare covariance matrices:
$$\begin{aligned} \operatorname{Cov}(\tilde\beta\mid X) &=\sigma^2AA^\mathsf{T}\\ &=\sigma^2(M_0+D)(M_0+D)^\mathsf{T}\\ &=\sigma^2M_0M_0^\mathsf{T}+\sigma^2DD^\mathsf{T}\\ &=\operatorname{Cov}(\hat\beta\mid X)+\sigma^2DD^\mathsf{T}. \end{aligned}$$The matrix $DD^\mathsf{T}$ is positive semidefinite because, for every vector $c$,
$$c^\mathsf{T}DD^\mathsf{T}c=\lVert D^\mathsf{T}c\rVert^2\ge 0.$$Therefore
$$\operatorname{Cov}(\tilde\beta\mid X)-\operatorname{Cov}(\hat\beta\mid X)\succeq 0.$$Or, stated one scalar target at a time,
$$\operatorname{Var}(c^\mathsf{T}\tilde\beta\mid X)-\operatorname{Var}(c^\mathsf{T}\hat\beta\mid X)=\sigma^2\lVert D^\mathsf{T}c\rVert^2\ge 0.$$That proves the theorem. Choosing $c=(1,0)^\mathsf{T}$ proves the result for the intercept in simple regression; choosing $c=(0,1)^\mathsf{T}$ proves it for the slope. No homework gap remains.
The same idea in the slope-only notation
The earlier scalar proof is still worth recognizing. Define
$$w_i=\frac{x_i-\bar x}{S_{xx}}, \qquad \hat\beta_1=\sum_{i=1}^n w_i y_i.$$The weights satisfy
$$\sum_i w_i=0, \qquad \sum_i w_i x_i=1.$$Any other linear unbiased slope estimator $\tilde\beta_1=\sum_i c_i y_i$ must satisfy the same two restrictions. Writing $k_i=c_i-w_i$ gives
$$\sum_i k_i=0, \qquad \sum_i k_ix_i=0,$$and therefore $\sum_iw_ik_i=0$. Hence
$$\operatorname{Var}(\tilde\beta_1\mid X)=\operatorname{Var}(\hat\beta_1\mid X)+\sigma^2\sum_i k_i^2\ge\operatorname{Var}(\hat\beta_1\mid X).$$This is the one-dimensional shadow of the matrix proof above.
Normality: useful, but not needed for BLUE
Nothing in the proof required $\varepsilon$ to be normally distributed. We used only conditional means, conditional covariances, full rank, and linear algebra.
If we add the stronger assumption
$$\varepsilon\mid X\sim N(0,\sigma^2I_n),$$then
$$\hat\beta\mid X\sim N\!\left(\beta,\sigma^2(X^\mathsf{T}X)^{-1}\right).$$When there are positive residual degrees of freedom, that normal model supports the familiar exact finite-sample $t$ and $F$ procedures once $\sigma^2$ is estimated in the usual way. It is an inference assumption, not a missing ingredient in the Gauss–Markov efficiency proof.
And under joint normality, zero covariance does imply independence. Without joint normality, calling uncorrelated errors “completely independent” is too strong.
What happens when an assumption fails?
- If $E(\varepsilon\mid X)\ne0$, OLS can be biased. Heteroskedasticity-robust standard errors do not repair that failure.
- If the errors are heteroskedastic or correlated but the conditional mean is still zero, OLS remains linear and unbiased, but the spherical-covariance BLUE claim no longer follows. The usual homoskedastic covariance formula is also wrong.
- If $X$ is rank-deficient, the coefficient vector is not uniquely identified by $(X^\mathsf{T}X)^{-1}X^\mathsf{T}y$ because the inverse does not exist.
- If the errors are nonnormal, Gauss–Markov can still hold. What changes is the route to exact small-sample inference, not the algebraic BLUE result.
So the right conclusion is not “OLS is always best.” It is:
Conditional on a full-rank design, when the errors have zero conditional mean and covariance $\sigma^2I$, OLS is the unique linear unbiased estimator with the smallest covariance matrix.
That is already a strong result. It is also exactly as strong as the assumptions—and no stronger.
References
- David A. Freedman, Notes on the Gauss–Markov theorem, University of California, Berkeley (2004). This gives the conditional assumptions and the full vector covariance proof.
- MIT OpenCourseWare, Regression Analysis, 18.S096, pp. 14–17. This states the uncorrelated, equal-variance assumptions and proves the minimum-variance result for arbitrary linear combinations of coefficients.
- MIT OpenCourseWare, Mathematical Statistics: Gaussian Linear Models, 18.655. This separates the Gaussian linear model and its distribution theory from the covariance-only Gauss–Markov result.
- National Institute of Standards and Technology, Linear Least Squares Regression. This is a practical reference for least-squares models that are linear in their unknown parameters.
Comments
Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.