Derivation of the Chi-Squared Distribution

Chi-squared is just a gamma in disguise — we prove Z² follows it with 1 degree of freedom and show how sample variance ties in before jumping to the t-distribution.

Faithful English translation of the author’s original Statistics #5 post. Original formula inconsistencies are preserved in source panels and explained separately below.

Alright, now it is time for the chi-squared distribution.

The chi-squared distribution is one of the gamma distributions with parameters α,λ\alpha,\lambda.

Among all those many gamma distributions, the one with α=n/2\alpha=n/2 and λ=1/2\lambda=1/2

is specifically called the “chi-squared distribution with nn degrees of freedom.”

The gamma distribution is

so if we substitute α=n/2\alpha=n/2 and λ=1/2\lambda=1/2 here,

Formula as printed in the source
Γ(α,λ)∼f(x)=xα−1e−λxλ−αΓ(α)\Gamma(\alpha,\lambda)\sim f(x)=\frac{x^{\alpha-1}e^{-\lambda x}}{\lambda^{-\alpha}\Gamma(\alpha)}
Preserved original image img_005.jpgPreserved original source image img_005.jpg

this is what we get.

Formula as printed in the source
χn2∼f(x)=xn/2−1e−x/2(1/2)−n/2Γ(n/2)=xn/2−1e−x/22n/2Γ(n/2)\chi_n^2\sim f(x)=\frac{x^{n/2-1}e^{-x/2}}{(1/2)^{-n/2}\Gamma(n/2)}=\frac{x^{n/2-1}e^{-x/2}}{2^{n/2}\Gamma(n/2)}
Preserved original image img_007.jpgPreserved original source image img_007.jpg

Since this is a gamma distribution, we can easily derive its moment-generating function, mean, variance, and so on.

Gamma and chi-squared: MGF, mean, variance

For the gamma distribution, the MGF, mean, and variance are respectively:

M(t)=(1−t/λ)−α,μ=α/λ,σ2=α/λ2M(t)=(1-t/\lambda)^{-\alpha},\qquad \mu=\alpha/\lambda,\qquad \sigma^2=\alpha/\lambda^2

Substitute α=n/2 and λ=1/2 here…

Chi-squared distribution:

M(t)=(1−t1/2)−n/2=(1−2t)−n/2M(t)=\left(1-\frac{t}{1/2}\right)^{-n/2}=(1-2t)^{-n/2}
μ=n/21/2=n,σ2=n/2(1/2)2=2n\mu=\frac{n/2}{1/2}=n,\qquad \sigma^2=\frac{n/2}{(1/2)^2}=2n
Preserved original image img_010.jpgPreserved original source image img_010.jpg

Why give this fellow its own definition when it is just one kind of gamma distribution? Because it is a little special, of course.

Yep, this one is special. It is connected to the normal distribution.

We usually write the random variable for the normal distribution as ZZ, right?

And here the distribution followed by ZZ is N(0,1)N(0,1)…

But look: the square of ZZ follows a chi-squared distribution with one degree of freedom!!!

χ2\chi^2
Preserved original image img_016.jpgPreserved original source image img_016.jpg

So, for a random variable Z∼N(0,1)Z\sim N(0,1), let us prove that its square follows the distribution shown here.

Z2∼χ12Z^2\sim\chi^2_1
Preserved original image img_019.jpgPreserved original source image img_019.jpg
CDF of the square of a standard normal

For the random variable Y=Z², let us calculate P[Y≤y]: the cumulative distribution function!

F(y)=P[Y≤y]=P[Z2≤y]=P[−y≤Z≤y]F(y)=P[Y\le y]=P[Z^2\le y]=P[-\sqrt y\le Z\le\sqrt y]

Writing it this way lets us work with the normal distribution.

To find the probability P[−√y≤Z≤√y], recall the normal probability density:

f(z)=12πe−z2/2f(z)=\frac{1}{\sqrt{2\pi}}e^{-z^2/2}
Preserved original image img_021.jpgPreserved original source image img_021.jpg
Derivation of the density
P[−y≤Z≤y]=∫−yye−z2/22π dz=2∫0ye−z2/22π dzP[-\sqrt y\le Z\le\sqrt y]=\int_{-\sqrt y}^{\sqrt y}\frac{e^{-z^2/2}}{\sqrt{2\pi}}\,dz=2\int_0^{\sqrt y}\frac{e^{-z^2/2}}{\sqrt{2\pi}}\,dz

The integrand is an even function.

z=x,dz=12x dxz=\sqrt x,\quad dz=\frac{1}{2\sqrt x}\,dx
=2∫0ye−x/22π12x dx=∫0ye−x/22πx dx=\textcolor{#b01828}{\cancel{\textcolor{#181818}{2}}}\int_0^y\frac{e^{-x/2}}{\sqrt{2\pi}}\frac{1}{\textcolor{#b01828}{\cancel{\textcolor{#181818}{2}}}\sqrt x}\,dx=\int_0^y\frac{e^{-x/2}}{\sqrt{2\pi x}}\,dx

The numerator factor 2 and denominator factor 2 cancel in this change-of-variable integrand.

F(y)=12π∫0yx1/2−1e−x/2 dxF(y)=\frac{1}{\sqrt{2\pi}}\int_0^y x^{1/2-1}e^{-x/2}\,dx

Differentiating both sides:

f(y)=12πy1/2−1e−y/2f(y)=\frac{1}{\sqrt{2\pi}}y^{1/2-1}e^{-y/2}

Recall the gamma function: √π=Γ(1/2).

f(y)=y1/2−1e−y/22 Γ(1/2)=y1/2−1e−y/221/2Γ(1/2)f(y)=\frac{y^{1/2-1}e^{-y/2}}{\sqrt2\,\Gamma(1/2)}=\frac{y^{1/2-1}e^{-y/2}}{2^{1/2}\Gamma(1/2)}

That is the chi-squared distribution with one degree of freedom.

Preserved original image img_022.jpgPreserved original source image img_022.jpg

Now that we have established that this square follows a chi-squared distribution (one kind of gamma distribution),

Z2Z^2
Preserved original image img_024.jpgPreserved original source image img_024.jpg

there is one more thing we can say for sure.

This sum follows a chi-squared distribution with nn—nn!!!!!!!!—degrees of freedom!!!!

Z12+Z22+Z32+⋯+Zn2Z_1^2+Z_2^2+Z_3^2+\cdots+Z_n^2
Preserved original image img_027.jpgPreserved original source image img_027.jpg
χn2\chi_n^2
Preserved original image img_029.jpgPreserved original source image img_029.jpg

Why?!?!?! A chi-squared distribution is still a gamma distribution,

and we already worked through the additivity of the gamma distribution in the previous post, remember??~~~~~~~~~~

(We will use this just below.)

I will skip this proof. Refer to the previous post; it is very easy to prove from that! Hehe.

Now let us put together just one more point about the chi-squared distribution.

I said it matters because it is related to the normal random variable ZZ,

but this fellow is also closely connected to samples.

And since that connection with samples will become a tool for getting to the tt distribution,

we should go through it before moving on.

We talked about the difference between population and sample waaaay!!!!~~~~~~ back,

so let us take another look.

First, let us check the notation again:

Population and sample notation

Population mean: μ. Sample mean:

Xˉ=1n∑Xi\bar X=\frac1n\sum X_i

Population variance: σ². Sample variance:

S2=1n−1∑(Xi−Xˉ)2S^2=\frac1{n-1}\sum(X_i-\bar X)^2
Preserved original image img_043.jpgPreserved original source image img_043.jpg

Once we bring samples into the discussion, the population mean and population variance become quantities to estimate… right…?

Let us briefly go through that connection:

Mean and variance of the sample mean
E(Xˉ)=E(1n∑Xi)=E(1n(X1+X2+⋯+Xn))\mathrm E(\bar X)=\mathrm E\left(\frac1n\sum X_i\right)=\mathrm E\left(\frac1n(X_1+X_2+\cdots+X_n)\right)
=1n(E(X1)+E(X2)+⋯+E(Xn))=1n(μ+μ+⋯+μ)=1nnμ=μ=\frac1n\left(\mathrm E(X_1)+\mathrm E(X_2)+\cdots+\mathrm E(X_n)\right)=\frac1n(\mu+\mu+\cdots+\mu)=\frac1n n\mu=\mu

E(Xᵢ)=μ is the random-sample assumption.

Var(Xˉ)=Var(1n(X1+X2+⋯+Xn))\mathrm{Var}(\bar X)=\mathrm{Var}\left(\frac1n(X_1+X_2+\cdots+X_n)\right)
=1n2(Var(X1)+Var(X2)+⋯+Var(Xn))=1n2(σ2+⋯+σ2)=1n2nσ2=σ2n=\frac1{n^2}\left(\mathrm{Var}(X_1)+\mathrm{Var}(X_2)+\cdots+\mathrm{Var}(X_n)\right)=\frac1{n^2}(\sigma^2+\cdots+\sigma^2)=\frac1{n^2}n\sigma^2=\frac{\sigma^2}{n}

This also uses the random-sample assumption.

Preserved original image img_046.jpgPreserved original source image img_046.jpg
Aside 1: decomposition of squared deviations
∑(Xi−μ)2=∑[(Xi−Xˉ)+(Xˉ−μ)]2\sum(X_i-\mu)^2=\sum[(X_i-\bar X)+(\bar X-\mu)]^2
=∑[(Xi−Xˉ)2+(Xˉ−μ)2+2(Xi−Xˉ)(Xˉ−μ)]=\sum[(X_i-\bar X)^2+(\bar X-\mu)^2+2(X_i-\bar X)(\bar X-\mu)]
=∑(Xi−Xˉ)2+∑(Xˉ−μ)2+2∑(Xi−Xˉ)(Xˉ−μ)=\sum(X_i-\bar X)^2+\sum(\bar X-\mu)^2+2\sum(X_i-\bar X)(\bar X-\mu)
=∑(Xi−Xˉ)2+∑(Xˉ−μ)2+2(Xˉ−μ)∑(Xi−Xˉ)=\sum(X_i-\bar X)^2+\sum(\bar X-\mu)^2+2(\bar X-\mu)\sum(X_i-\bar X)

The last sum is zero.

=∑(Xi−Xˉ)2+n(Xˉ−μ)2=\sum(X_i-\bar X)^2+n(\bar X-\mu)^2
∴ ∑(Xi−μ)2=∑(Xi−Xˉ)2+n(Xˉ−μ)2\therefore\ \sum(X_i-\mu)^2=\sum(X_i-\bar X)^2+n(\bar X-\mu)^2
Preserved original image img_047.jpgPreserved original source image img_047.jpg
Expectation of the sample variance
E(S2)=E(1n−1∑(Xi−Xˉ)2)\mathrm E(S^2)=\mathrm E\left(\frac1{\textcolor{#b01828}{n-1}}\sum(X_i-\bar X)^2\right)

The asterisk emphasizes the divisor n−1 in the expectation of S².

=1n−1E(∑(Xi−μ)2−n(Xˉ−μ)2)=\frac1{\textcolor{#b01828}{n-1}}\mathrm E\left(\sum(X_i-\mu)^2-n(\bar X-\mu)^2\right)
=1n−1[∑E[(Xi−μ)2]−nE[(Xˉ−μ)2]]=\frac1{n-1}\left[\sum\mathrm E[(X_i-\mu)^2]-n\mathrm E[(\bar X-\mu)^2]\right]

Each sample observation has variance σ² (=Var(Xᵢ)).

The sample mean has variance σ²/n (=Var(X̄)).

=1n−1[∑σ2−nσ2n]=1n−1(nσ2−σ2)=1n−1(n−1)σ2=σ2=\frac1{n-1}\left[\sum\sigma^2-n\frac{\sigma^2}{n}\right]=\frac1{n-1}(n\sigma^2-\sigma^2)=\frac1{n-1}(n-1)\sigma^2=\sigma^2
∴ E(S2)=σ2\therefore\ \mathrm E(S^2)=\sigma^2
Preserved original image img_048.jpgPreserved original source image img_048.jpg

When calculating the sample variance shown here, we divide by n−1n-1 rather than nn

S2S^2
Preserved original image img_050.jpgPreserved original source image img_050.jpg

because, as you probably know, we reduce the degrees of freedom by one.

But we have also discovered something mathematically!!!!!

The purpose was to make the expectation of this sample variance unbiased.

S2S^2
Preserved original image img_054.jpgPreserved original source image img_054.jpg

(Well, reducing the degrees of freedom by one is what makes it unbiased, so these are really the same point… hehe.)

But what we were trying to do here

was not that;

we wanted to say that the chi-squared distribution is related to this sample variance…

S2S^2
Preserved original image img_060.jpgPreserved original source image img_060.jpg

Anyway, back to where we were!!!!!

Suppose the population follows the Gaussian distribution shown here,

N(μ,σ2)N(\mu,\sigma^2)
Preserved original image img_064.jpgPreserved original source image img_064.jpg

and we take a random sample X1,X2,X3,…,XnX_1,X_2,X_3,\ldots,X_n of size nn from it.

If its sample variance is the quantity shown here,

S2S^2
Preserved original image img_068.jpgPreserved original source image img_068.jpg

“The statistic shown here follows a chi-squared distribution with n−1n-1 degrees of freedom.”

(n−1)S2σ2\frac{(n-1)S^2}{\sigma^2}
Preserved original image img_071.jpgPreserved original source image img_071.jpg

That is what we were trying to prove!!!!!!!!!!!!!

Let us gooooooooooo!

Relating sample variance to chi-squared variables
S2=1n−1∑(Xi−Xˉ)2S^2=\frac1{n-1}\sum(X_i-\bar X)^2
(n−1)S2=∑(Xi−Xˉ)2(n-1)S^2=\sum(X_i-\bar X)^2

Divide both sides by σ².

(n−1)S2σ2=∑(Xi−Xˉ)2σ2\frac{(n-1)S^2}{\sigma^2}=\frac{\sum(X_i-\bar X)^2}{\sigma^2}

Use Aside 1 above; this is why we established it.

=∑(Xi−μ)2−n(Xˉ−μ)2σ2=\frac{\sum(X_i-\mu)^2-n(\bar X-\mu)^2}{\sigma^2}
=∑(Xi−μ)2σ2−n(Xˉ−μ)2σ2=∑(Xi−μ)2σ2−n(Xˉ−μσ)2=\frac{\sum(X_i-\mu)^2}{\sigma^2}-\frac{n(\bar X-\mu)^2}{\sigma^2}=\frac{\sum(X_i-\mu)^2}{\sigma^2}-n\left(\frac{\bar X-\mu}{\sigma}\right)^2
=∑(Xi−μ)2σ2−(Xˉ−μσ2/n)2=\frac{\sum(X_i-\mu)^2}{\sigma^2}-\left(\frac{\bar X-\mu}{\sqrt{\sigma^2/n}}\right)^2

Arrange it this way.

First term: chi-squared with n degrees of freedom. Second term: chi-squared with one degree of freedom.

Preserved original image img_075.jpgPreserved original source image img_075.jpg
MGF argument as written in the source
(n−1)S2σ2=∑(Xi−Xˉ)2σ2\frac{(n-1)S^2}{\sigma^2}=\frac{\sum(X_i-\bar X)^2}{\sigma^2}
∑(Xi−Xˉ)2σ2=∑(Xi−μ)2σ2−(Xˉ−μσ2/n)2\frac{\sum(X_i-\bar X)^2}{\sigma^2}=\frac{\sum(X_i-\mu)^2}{\sigma^2}-\left(\frac{\bar X-\mu}{\sqrt{\sigma^2/n}}\right)^2

The terms on the right follow χ²ₙ and χ²₁.

Rewrite it as

∑(Xi−μ)2σ2⏟W=∑(Xi−Xˉ)2σ2⏟U+(Xˉ−μσ2/n)2⏟V\underbrace{\frac{\sum(X_i-\mu)^2}{\sigma^2}}_W=\underbrace{\frac{\sum(X_i-\bar X)^2}{\sigma^2}}_U+\underbrace{\left(\frac{\bar X-\mu}{\sqrt{\sigma^2/n}}\right)^2}_V

Then the relationship between moment-generating functions is

Editorial clarification before using the MGF identity: for an independent normal sample, the sample mean and residual vector are jointly normal with zero covariance, hence independent. Thus U and V are independent, and the following product identity is valid.

MW(t)=MU(t)MV(t)M_W(t)=M_U(t)M_V(t)

The next symbolic ratio is retained exactly as written; it contradicts the preceding line.

MU(t)=MV(t)MW(t)M_U(t)=\frac{M_V(t)}{M_W(t)}

The source says we know the distributions of W and V, then writes:

=(1−2t)−n/2(1−2t)−1/2=(1−2t)−(n−1)/2=\frac{(1-2t)^{-n/2}}{(1-2t)^{-1/2}}=(1-2t)^{-(n-1)/2}

Therefore U, namely the following statistic, follows χ² with n−1 degrees of freedom!

U=∑(Xi−Xˉ)2σ2=(n−1)S2σ2∼χn−12U=\frac{\sum(X_i-\bar X)^2}{\sigma^2}=\frac{(n-1)S^2}{\sigma^2}\sim\chi^2_{n-1}
Preserved original image img_076.jpgPreserved original source image img_076.jpg

Done~

↓ Once more, neatly written up on the iPad!

iPad recap 1: chi-squared distribution

Chi-Squared Distribution

The chi-squared distribution is one of the gamma distributions! The gamma distribution with α=n/2 and λ=1/2 is called chi-squared with n degrees of freedom.

The source writes the gamma density with rate λ below.

Γ(α,λ)∼f(x)=xα−1e−λxλ−αΓ(α)\Gamma(\alpha,\lambda)\sim f(x)=\frac{x^{\alpha-1}e^{-\lambda x}}{\lambda^{-\alpha}\Gamma(\alpha)}

Meaning: when events occur at rate λ per unit time, observe until the αth event; the probability associated with time x is written this way.

Substitute α=n/2 and λ=1/2:

χn2∼f(x)=xn/2−1e−x/22n/2Γ(n/2)\chi_n^2\sim f(x)=\frac{x^{n/2-1}e^{-x/2}}{2^{n/2}\Gamma(n/2)}

The source describes waiting at rate 1/2 for event n/2; the probability associated with time x is written this way. Substituting gives this; what does it mean?

M(t)=(1−2t)−n/2,μ=n,σ2=2nM(t)=(1-2t)^{-n/2},\quad\mu=n,\quad\sigma^2=2n

The chi-squared distribution is the distribution of Z² when Z follows N(0,1)! Shall we check?

Y=Z2,Z∼N(0,1)Y=Z^2,\quad Z\sim N(0,1)

Calculate P[Y≤y] (the CDF):

F(y)=P[Y≤y]=P[Z2≤y]=P[−y≤Z≤y]F(y)=P[Y\le y]=P[Z^2\le y]=P[-\sqrt y\le Z\le\sqrt y]

The normal density is:

f(z)=12πe−z2/2f(z)=\frac1{\sqrt{2\pi}}e^{-z^2/2}
F(y)=∫−yye−z2/22π dz=2∫0ye−z2/22π dzF(y)=\int_{-\sqrt y}^{\sqrt y}\frac{e^{-z^2/2}}{\sqrt{2\pi}}\,dz=2\int_0^{\sqrt y}\frac{e^{-z^2/2}}{\sqrt{2\pi}}\,dz

Integrating an even function.

z=x,dz=12x dxz=\sqrt x,\quad dz=\frac1{2\sqrt x}\,dx
=2∫0y12πe−x/212x dx=\textcolor{#087522}{\cancel{\textcolor{#181818}{2}}}\int_0^y\frac{1}{\sqrt{2\pi}}e^{-x/2}\frac{1}{\textcolor{#087522}{\cancel{\textcolor{#181818}{2}}}\sqrt{x}}\,dx

The factors of 2 cancel.

F(y)=12π∫0yx−1/2e−x/2 dxF(y)=\frac1{\sqrt{2\pi}}\int_0^y x^{-1/2}e^{-x/2}\,dx

Differentiate both sides to obtain the probability density!

f(y)=12πy1/2−1e−y/2=y1/2−1e−y/221/2Γ(1/2)f(y)=\frac1{\sqrt{2\pi}}y^{1/2-1}e^{-y/2}=\frac{y^{1/2-1}e^{-y/2}}{2^{1/2}\Gamma(1/2)}
π=Γ(1/2)\sqrt\pi=\Gamma(1/2)

You can check this using the gamma function.

The density of Z² really is chi-squared with one degree of freedom… wow…

Preserved original image img_079.pngPreserved original source image img_079.png
iPad recap 2: samples and unbiased variance

We checked that chi-squared is one kind of gamma distribution. Earlier we also checked gamma additivity, right?

Z2∼χ12,Z12+⋯+Zn2∼χn2Z^2\sim\chi^2_1,\qquad Z_1^2+\cdots+Z_n^2\sim\chi^2_n

There are n terms. This follows from gamma additivity, so let us move on.

One more thing! Chi-squared is connected with sample variance. Let us check:

(n−1)S2σ2∼χn−12\frac{(n-1)S^2}{\sigma^2}\sim\chi^2_{n-1}

First put the notation together: population mean μ; population variance σ²; sample mean and variance:

Xˉ=1n∑Xi,S2=1n−1∑(Xi−Xˉ)2\bar X=\frac1n\sum X_i,\qquad S^2=\frac1{\textcolor{#1546b9}{n-1}}\sum(X_i-\bar X)^2

We will also explain why the divisor is n−1.

Relationship between population and sample; expectation of the sample mean:

E(Xˉ)=E(1n∑Xi)=E(1n(X1+X2+⋯+Xn))\mathrm E(\bar X)=\mathrm E\left(\frac1n\sum X_i\right)=\mathrm E\left(\frac1n(X_1+X_2+\cdots+X_n)\right)
=1n(E(X1)+E(X2)+⋯+E(Xn))=1n(μ+μ+⋯+μ)=1nnμ=μ=\frac1n\left(\mathrm E(X_1)+\mathrm E(X_2)+\cdots+\mathrm E(X_n)\right)=\frac1n(\mu+\mu+\cdots+\mu)=\frac1n n\mu=\mu

E(Xᵢ)=μ is the random-sample assumption.

Var(Xˉ)=Var(1n(X1+X2+⋯+Xn))\mathrm{Var}(\bar X)=\mathrm{Var}\left(\frac1n(X_1+X_2+\cdots+X_n)\right)
=1n2(Var(X1)+Var(X2)+⋯+Var(Xn))=1n2(σ2+⋯+σ2)=1n2nσ2=σ2n=\frac1{n^2}\left(\mathrm{Var}(X_1)+\mathrm{Var}(X_2)+\cdots+\mathrm{Var}(X_n)\right)=\frac1{n^2}(\sigma^2+\cdots+\sigma^2)=\frac1{n^2}n\sigma^2=\frac{\sigma^2}{n}

This also uses the random-sample assumption: Var(Xᵢ)=σ².

Aside: expansion and cancellation of the cross term.

∑(Xi−μ)2=∑[(Xi−Xˉ)+(Xˉ−μ)]2\sum(X_i-\mu)^2=\sum[(X_i\textcolor{#b01828}{-\bar X})+(\bar X-\mu)]^2
=∑[(Xi−Xˉ)2+(Xˉ−μ)2+2(Xi−Xˉ)(Xˉ−μ)]=\sum[(X_i-\bar X)^2+(\bar X-\mu)^2+2(X_i-\bar X)(\bar X-\mu)]
=∑(Xi−Xˉ)2+∑(Xˉ−μ)2+2∑(Xi−Xˉ)(Xˉ−μ)=\sum(X_i-\bar X)^2+\sum(\bar X-\mu)^2+2\sum(X_i-\bar X)(\bar X-\mu)
=∑(Xi−Xˉ)2+∑(Xˉ−μ)2+2(Xˉ−μ)∑(Xi−Xˉ)=\sum(X_i-\bar X)^2+\sum(\bar X-\mu)^2+2(\bar X-\mu)\sum(X_i-\bar X)
∑(Xi−Xˉ)=0\sum(X_i-\bar X)=0
=∑(Xi−Xˉ)2+n(Xˉ−μ)2=\sum(X_i-\bar X)^2+n(\bar X-\mu)^2
∴ ∑(Xi−μ)2=∑(Xi−Xˉ)2+n(Xˉ−μ)2\therefore\ \sum(X_i-\mu)^2=\sum(X_i-\bar X)^2+n(\bar X-\mu)^2
∑(Xi−Xˉ)2=∑(Xi−μ)2−n(Xˉ−μ)2\sum(X_i-\bar X)^2=\sum(X_i-\mu)^2-n(\bar X-\mu)^2

Let us look at the significance of n−1!

E(S2)=E(1n−1∑(Xi−Xˉ)2)\mathrm E(S^2)=\mathrm E\left(\frac1{\textcolor{#b01280}{n-1}}\sum(X_i-\bar X)^2\right)
=1n−1E(∑(Xi−μ)2−n(Xˉ−μ)2)=\frac1{n-1}\mathrm E\left(\textcolor{#b01828}{\sum(X_i-\mu)^2-n(\bar X-\mu)^2}\right)
=1n−1[∑E[(Xi−μ)2]−nE[(Xˉ−μ)2]]=\frac1{n-1}\left[\sum\mathrm E[(X_i-\mu)^2]-n\mathrm E[(\bar X-\mu)^2]\right]

Each sample observation has variance σ² (=Var(Xᵢ)).

The sample mean has variance σ²/n (=Var(X̄)).

=1n−1[∑σ2−nσ2n]=1n−1(nσ2−σ2)=1n−1(n−1)σ2=σ2=\frac1{n-1}\left[\sum\sigma^2-n\frac{\sigma^2}{n}\right]=\frac1{n-1}(n\sigma^2-\sigma^2)=\frac1{n-1}(n-1)\sigma^2=\sigma^2
∴ E(S2)=σ2\therefore\ \mathrm E(S^2)=\sigma^2

We must divide by this number to make the expectation of sample variance unbiased. (This is really the same point as before…)

Preserved original image img_080.pngPreserved original source image img_080.png
iPad recap 3: sample-variance distribution

Finally, let us check the following fact!

(n−1)S2σ2∼χn−12\frac{(n-1)S^2}{\sigma^2}\sim\chi^2_{n-1}
S2=1n−1∑(Xi−Xˉ)2S^2=\frac1{n-1}\sum(X_i-\bar X)^2
(n−1)S2=∑(Xi−Xˉ)2(n-1)S^2=\sum(X_i-\bar X)^2

Divide both sides by σ².

(n−1)S2σ2=∑(Xi−Xˉ)2σ2\frac{(n-1)S^2}{\sigma^2}=\frac{\sum(X_i-\bar X)^2}{\sigma^2}

Use the red rearranged identity in Aside 1 for the red numerator in the following fraction.

=∑(Xi−μ)2−n(Xˉ−μ)2σ2=\frac{\textcolor{#b01828}{\sum(X_i-\mu)^2-n(\bar X-\mu)^2}}{\sigma^2}
=∑(Xi−μ)2σ2−n(Xˉ−μ)2σ2=∑(Xi−μ)2σ2−n(Xˉ−μσ)2=\frac{\sum(X_i-\mu)^2}{\sigma^2}-\frac{n(\bar X-\mu)^2}{\sigma^2}=\frac{\sum(X_i-\mu)^2}{\sigma^2}-n\left(\frac{\bar X-\mu}{\sigma}\right)^2
=∑(Xi−μ)2σ2−(Xˉ−μσ2/n)2=\frac{\sum(X_i-\mu)^2}{\sigma^2}-\left(\frac{\bar X-\mu}{\sqrt{\sigma^2/n}}\right)^2

Arrange it this way.

First term: chi-squared with n degrees of freedom. Second term: chi-squared with one degree of freedom.

Aside: the expansion used here.

∑(Xi−μ)2=∑[(Xi−Xˉ)+(Xˉ−μ)]2\sum(X_i-\mu)^2=\sum[(X_i\textcolor{#b01828}{-\bar X})+(\bar X-\mu)]^2
=∑[(Xi−Xˉ)2+(Xˉ−μ)2+2(Xi−Xˉ)(Xˉ−μ)]=\sum[(X_i-\bar X)^2+(\bar X-\mu)^2+2(X_i-\bar X)(\bar X-\mu)]
=∑(Xi−Xˉ)2+∑(Xˉ−μ)2+2∑(Xi−Xˉ)(Xˉ−μ)=\sum(X_i-\bar X)^2+\sum(\bar X-\mu)^2+2\sum(X_i-\bar X)(\bar X-\mu)
=∑(Xi−Xˉ)2+∑(Xˉ−μ)2+2(Xˉ−μ)∑(Xi−Xˉ)=\sum(X_i-\bar X)^2+\sum(\bar X-\mu)^2+2(\bar X-\mu)\sum(X_i-\bar X)
∑(Xi−Xˉ)=0\sum(X_i-\bar X)=0
=∑(Xi−Xˉ)2+n(Xˉ−μ)2=\sum(X_i-\bar X)^2+n(\bar X-\mu)^2
∴ ∑(Xi−μ)2=∑(Xi−Xˉ)2+n(Xˉ−μ)2\therefore\ \sum(X_i-\mu)^2=\sum(X_i-\bar X)^2+n(\bar X-\mu)^2

Rearrange the expansion to obtain Σ(Xᵢ−X̄)²=Σ(Xᵢ−μ)²−n(X̄−μ)².

There are n squared Z terms, following χ²ₙ, and one squared Z term, following χ²₁.

Putting the preceding work together:

(n−1)S2σ2=∑(Xi−Xˉ)2σ2\frac{(n-1)S^2}{\sigma^2}=\frac{\sum(X_i-\bar X)^2}{\sigma^2}
∑(Xi−Xˉ)2σ2=∑(Xi−μ)2σ2−(Xˉ−μσ2/n)2\frac{\sum(X_i-\bar X)^2}{\sigma^2}=\frac{\sum(X_i-\mu)^2}{\sigma^2}-\left(\frac{\bar X-\mu}{\sqrt{\sigma^2/n}}\right)^2

The terms on the right follow χ²ₙ and χ²₁.

Rewrite it as

∑(Xi−μ)2σ2⏟W=∑(Xi−Xˉ)2σ2⏟U+(Xˉ−μσ2/n)2⏟V\underbrace{\frac{\sum(X_i-\mu)^2}{\sigma^2}}_{\textcolor{#1546b9}{W}}=\underbrace{\frac{\sum(X_i-\bar X)^2}{\sigma^2}}_{\textcolor{#1546b9}{U}}+\underbrace{\left(\frac{\bar X-\mu}{\sqrt{\sigma^2/n}}\right)^2}_{\textcolor{#1546b9}{V}}

Write these terms as W, U, V. Then the MGF relationship is as shown.

Then the relationship between moment-generating functions is

Editorial clarification before using the MGF identity: for an independent normal sample, the sample mean and residual vector are jointly normal with zero covariance, hence independent. Thus U and V are independent, and the following product identity is valid.

MW(t)=MU(t)MV(t)M_W(t)=M_U(t)M_V(t)

The next symbolic ratio is retained exactly as written; it contradicts the preceding line.

MU(t)=MV(t)MW(t)M_U(t)=\frac{M_V(t)}{M_W(t)}

The source says we know the distributions of W and V, then writes:

=(1−2t)−n/2(1−2t)−1/2=(1−2t)−(n−1)/2=\frac{(1-2t)^{-n/2}}{(1-2t)^{-1/2}}=(1-2t)^{-(n-1)/2}

Therefore U, namely the following statistic, follows χ² with n−1 degrees of freedom!

U=∑(Xi−Xˉ)2σ2=(n−1)S2σ2∼χn−12U=\frac{\sum(X_i-\bar X)^2}{\sigma^2}=\frac{(n-1)S^2}{\sigma^2}\sim\chi^2_{n-1}
Preserved original image img_081.pngPreserved original source image img_081.png

P.S. (historical author note)

Did you know?

“I have converted all the blog posts to PDF,

and I am selling the PDF materials :-)”

Historical PDF sales post

Historical link preview: “Blog-post PDFs (ver. 2.0) for sale (the physics and finance I studied). Purchase instructions are below ~ Hello! If there are parts of the blog posts you are dissatisfied with, too much…” — blog.naver.com. The source excerpt ends here. This archived note records the author’s past announcement; current sale or availability is not asserted.

Original plotting code (14 lines). Only nonbreaking spaces have been normalized to ordinary spaces. The gray numbers 1–14 are line-number UI; “cs” identifies the ColorScripter widget.

n = [1., 2., 3., 4., 5.]
for i in n:
    alpha = i / 2           # i = n
    llambda = 0.5
    scale= 1./llambda
    x = np.linspace(0, 10, 100)
    y = sc.gamma.pdf(x, a=alpha, scale=scale)
    plt.plot(x, y, linewidth=2.0, label = 'n=%s' % i)
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.title(r'$\chi$-squared Distribution')
plt.savefig('4.Chi-squared Distribution.jpeg')

Source widget attribution: cs / ColorScripter.

Chi-squared density curves for n=1, 2, 3, 4, and 5
Original plot: chi-squared density curves. The source title uses χ (visually similar to x); values and curves are retained unchanged.

Comments

Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.