Derivation of the Chi-Squared DistributionPhysicsBasic Statistics I Studied5Chi-squared is just a gamma in disguise — we prove Z² follows it with 1 degree of freedom and show how sample variance ties in before jumping to the t-distribution.
Chi-squared is just a gamma in disguise — we prove Z² follows it with 1 degree of freedom and show how sample variance ties in before jumping to the t-distribution.
·11 min read·gdpark
Faithful English translation of the author’s original Statistics #5 post. Original formula inconsistencies are preserved in source panels and explained separately below.
Alright, now it is time for the chi-squared distribution.
The chi-squared distribution is one of the gamma distributions with parameters .
Among all those many gamma distributions, the one with and
is specifically called the “chi-squared distribution with degrees of freedom.”
The gamma distribution is
so if we substitute and here,
Formula as printed in the source
Preserved original image img_005.jpg
this is what we get.
Formula as printed in the source
Preserved original image img_007.jpg
Since this is a gamma distribution, we can easily derive its moment-generating function, mean, variance, and so on.
Gamma and chi-squared: MGF, mean, variance
For the gamma distribution, the MGF, mean, and variance are respectively:
Substitute α=n/2 and λ=1/2 here…
Chi-squared distribution:
Preserved original image img_010.jpg
Why give this fellow its own definition when it is just one kind of gamma distribution? Because it is a little special, of course.
Yep, this one is special. It is connected to the normal distribution.
We usually write the random variable for the normal distribution as , right?
And here the distribution followed by is …
But look: the square of follows a chi-squared distribution with one degree of freedom!!!
Preserved original image img_016.jpg
So, for a random variable , let us prove that its square follows the distribution shown here.
Preserved original image img_019.jpgCDF of the square of a standard normal
For the random variable Y=Z², let us calculate P[Y≤y]: the cumulative distribution function!
Writing it this way lets us work with the normal distribution.
To find the probability P[−√y≤Z≤√y], recall the normal probability density:
Preserved original image img_021.jpgDerivation of the density
The integrand is an even function.
=2∫0y2πe−x/22x1dx=∫0y2πxe−x/2dx
The numerator factor 2 and denominator factor 2 cancel in this change-of-variable integrand.
Differentiating both sides:
Recall the gamma function: √π=Γ(1/2).
That is the chi-squared distribution with one degree of freedom.
Preserved original image img_022.jpg
Now that we have established that this square follows a chi-squared distribution (one kind of gamma distribution),
Preserved original image img_024.jpg
there is one more thing we can say for sure.
This sum follows a chi-squared distribution with —!!!!!!!!—degrees of freedom!!!!
Preserved original image img_027.jpg
Preserved original image img_029.jpg
Why?!?!?! A chi-squared distribution is still a gamma distribution,
and we already worked through the additivity of the gamma distribution in the previous post, remember??~~~~~~~~~~
(We will use this just below.)
I will skip this proof. Refer to the previous post; it is very easy to prove from that! Hehe.
Now let us put together just one more point about the chi-squared distribution.
I said it matters because it is related to the normal random variable ,
but this fellow is also closely connected to samples.
And since that connection with samples will become a tool for getting to the distribution,
we should go through it before moving on.
We talked about the difference between population and sample waaaay!!!!~~~~~~ back,
so let us take another look.
First, let us check the notation again:
Population and sample notation
Population mean: μ. Sample mean:
Population variance: σ². Sample variance:
Preserved original image img_043.jpg
Once we bring samples into the discussion, the population mean and population variance become quantities to estimate… right…?
Let us briefly go through that connection:
Mean and variance of the sample mean
E(Xᵢ)=μ is the random-sample assumption.
This also uses the random-sample assumption.
Preserved original image img_046.jpgAside 1: decomposition of squared deviations
The last sum is zero.
Preserved original image img_047.jpgExpectation of the sample variance
The asterisk emphasizes the divisor n−1 in the expectation of S².
Each sample observation has variance σ² (=Var(Xᵢ)).
The sample mean has variance σ²/n (=Var(X̄)).
Preserved original image img_048.jpg
When calculating the sample variance shown here, we divide by rather than
Preserved original image img_050.jpg
because, as you probably know, we reduce the degrees of freedom by one.
But we have also discovered something mathematically!!!!!
The purpose was to make the expectation of this sample variance unbiased.
Preserved original image img_054.jpg
(Well, reducing the degrees of freedom by one is what makes it unbiased, so these are really the same point… hehe.)
But what we were trying to do here
was not that;
we wanted to say that the chi-squared distribution is related to this sample variance…
Preserved original image img_060.jpg
Anyway, back to where we were!!!!!
Suppose the population follows the Gaussian distribution shown here,
Preserved original image img_064.jpg
and we take a random sample of size from it.
If its sample variance is the quantity shown here,
Preserved original image img_068.jpg
“The statistic shown here follows a chi-squared distribution with degrees of freedom.”
Preserved original image img_071.jpg
That is what we were trying to prove!!!!!!!!!!!!!
Let us gooooooooooo!
Relating sample variance to chi-squared variables
Divide both sides by σ².
Use Aside 1 above; this is why we established it.
Arrange it this way.
First term: chi-squared with n degrees of freedom. Second term: chi-squared with one degree of freedom.
Preserved original image img_075.jpgMGF argument as written in the source
The terms on the right follow χ²ₙ and χ²₁.
Rewrite it as
Then the relationship between moment-generating functions is
Editorial clarification before using the MGF identity: for an independent normal sample, the sample mean and residual vector are jointly normal with zero covariance, hence independent. Thus U and V are independent, and the following product identity is valid.
The next symbolic ratio is retained exactly as written; it contradicts the preceding line.
The source says we know the distributions of W and V, then writes:
Therefore U, namely the following statistic, follows χ² with n−1 degrees of freedom!
Preserved original image img_076.jpg
Done~
↓ Once more, neatly written up on the iPad!
iPad recap 1: chi-squared distribution
Chi-Squared Distribution
The chi-squared distribution is one of the gamma distributions! The gamma distribution with α=n/2 and λ=1/2 is called chi-squared with n degrees of freedom.
The source writes the gamma density with rate λ below.
Meaning: when events occur at rate λ per unit time, observe until the αth event; the probability associated with time x is written this way.
Substitute α=n/2 and λ=1/2:
The source describes waiting at rate 1/2 for event n/2; the probability associated with time x is written this way. Substituting gives this; what does it mean?
The chi-squared distribution is the distribution of Z² when Z follows N(0,1)! Shall we check?
Calculate P[Y≤y] (the CDF):
The normal density is:
Integrating an even function.
=2∫0y2π1e−x/22x1dx
The factors of 2 cancel.
Differentiate both sides to obtain the probability density!
You can check this using the gamma function.
The density of Z² really is chi-squared with one degree of freedom… wow…
Preserved original image img_079.pngiPad recap 2: samples and unbiased variance
We checked that chi-squared is one kind of gamma distribution. Earlier we also checked gamma additivity, right?
There are n terms. This follows from gamma additivity, so let us move on.
One more thing! Chi-squared is connected with sample variance. Let us check:
First put the notation together: population mean μ; population variance σ²; sample mean and variance:
We will also explain why the divisor is n−1.
Relationship between population and sample; expectation of the sample mean:
E(Xᵢ)=μ is the random-sample assumption.
This also uses the random-sample assumption: Var(Xᵢ)=σ².
Aside: expansion and cancellation of the cross term.
Let us look at the significance of n−1!
Each sample observation has variance σ² (=Var(Xᵢ)).
The sample mean has variance σ²/n (=Var(X̄)).
We must divide by this number to make the expectation of sample variance unbiased. (This is really the same point as before…)
Preserved original image img_080.pngiPad recap 3: sample-variance distribution
Finally, let us check the following fact!
Divide both sides by σ².
Use the red rearranged identity in Aside 1 for the red numerator in the following fraction.
Arrange it this way.
First term: chi-squared with n degrees of freedom. Second term: chi-squared with one degree of freedom.
Aside: the expansion used here.
Rearrange the expansion to obtain Σ(Xᵢ−X̄)²=Σ(Xᵢ−μ)²−n(X̄−μ)².
There are n squared Z terms, following χ²ₙ, and one squared Z term, following χ²₁.
Putting the preceding work together:
The terms on the right follow χ²ₙ and χ²₁.
Rewrite it as
Write these terms as W, U, V. Then the MGF relationship is as shown.
Then the relationship between moment-generating functions is
Editorial clarification before using the MGF identity: for an independent normal sample, the sample mean and residual vector are jointly normal with zero covariance, hence independent. Thus U and V are independent, and the following product identity is valid.
The next symbolic ratio is retained exactly as written; it contradicts the preceding line.
The source says we know the distributions of W and V, then writes:
Therefore U, namely the following statistic, follows χ² with n−1 degrees of freedom!
Original plotting code (14 lines). Only nonbreaking spaces have been normalized to ordinary spaces. The gray numbers 1–14 are line-number UI; “cs” identifies the ColorScripter widget.
n = [1., 2., 3., 4., 5.]
for i in n:
alpha = i / 2 # i = n
llambda = 0.5
scale= 1./llambda
x = np.linspace(0, 10, 100)
y = sc.gamma.pdf(x, a=alpha, scale=scale)
plt.plot(x, y, linewidth=2.0, label = 'n=%s' % i)
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.title(r'$\chi$-squared Distribution')
plt.savefig('4.Chi-squared Distribution.jpeg')
Comments
Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.