Derivation of the Student's t-Distribution

Where the 'Student' name came from, why we ditch the Z-stat when σ is unknown, and a full derivation of the t-distribution PDF — plus properties and a worked example.

English translation of the original Statistics #6 post. Source notation and the author’s original claims are preserved. Separate explanatory notes identify and correct source errors.

First, why is this distribution called the Student t-distribution?

W. S. Grosset, who worked at a brewery, first discovered and used it while working there.

Apparently, he did not want other breweries to use it in the same way, so he published it under the pen name “student t.”

First, let us look at what the t-distribution is used for.

To estimate the population mean μ, we use the sample mean Sample mean X bar. The distribution related to Sample mean X bar repeated is

and we use Z statistic.

When we do not know Population variance sigma squared, we cannot use the Z statistic,

so, as an alternative, we use T statistic, the T statistic.

The distribution followed by this T statistic is called the “t-distribution.”

The probability density function of this t-distribution is

Student t density, apparently.....

We cannot just skip past this without deriving it, can we????????

Actually, the derivation is the whole post, lol lol lol lol hehe.

Let us get into the derivation!

Rewriting T and introducing Z and U

First, let us change the form of T a little.

Xˉ−μS/n=Xˉ−μσ/n⋅σS\frac{\bar X-\mu}{S/\sqrt n}=\frac{\bar X-\mu}{\sigma/\sqrt n}\cdot\frac{\sigma}{S}
=Xˉ−μσ/n⋅1S/σ=\frac{\bar X-\mu}{\sigma/\sqrt n}\cdot\frac{1}{S/\sigma}
=Xˉ−μσ/n⋅1nS2σ2⋅1n=\frac{\bar X-\mu}{\sigma/\sqrt n}\cdot\frac{1}{\sqrt{\frac{nS^2}{\sigma^2}\cdot\frac1n}}
=Xˉ−μσ/nnS2/σ2n=\frac{\frac{\bar X-\mu}{\sigma/\sqrt n}}{\sqrt{\frac{nS^2/\sigma^2}{n}}}

Phew...

After rewriting it this way, the numerator is

Xˉ−μσ/n=Z∼N(0,1)\frac{\bar X-\mu}{\sigma/\sqrt n}=Z\sim N(0,1)

and we saw the denominator’s nS²/σ² when discussing the chi-squared distribution!

There, we proved that (n−1)S/σ² follows χ² with n−1 degrees of freedom, so we can say that nS²/σ² follows χ² with n degrees of freedom.

(n−1)Sσ2∼χn−12,nS2σ2∼χn2\frac{(n-1)S}{\sigma^2}\sim\chi^2_{n-1},\qquad\frac{nS^2}{\sigma^2}\sim\chi^2_n

Then, for that T, substitute

Z=Xˉ−μσ/n,U=nS2σ2Z=\frac{\bar X-\mu}{\sigma/\sqrt n},\qquad U=\frac{nS^2}{\sigma^2}
Z∼N(0,1),U∼χn2Z\sim N(0,1),\qquad U\sim\chi^2_n
Preserved original img_024.pngPreserved original Korean source img_024.png
Joint density and cumulative distribution

As we have always done, let us calculate the cumulative distribution of this guy:

T=ZU/nT=\frac{Z}{\sqrt{U/n}}

Since Z and U are statistically independent,

their joint probability density function (joint probability density function) is

If “statistically independent” or “joint probability density function” is unfamiliar, check the mathematical definition and have a think... It is not difficult.

g(z,u)=12πe−z2/2⋅un/2−1e−u/22n/2Γ(n/2)g(z,u)=\frac{1}{\sqrt{2\pi}}e^{-z^2/2}\cdot\frac{u^{n/2-1}e^{-u/2}}{2^{n/2}\Gamma(n/2)}

Now, to calculate T’s cumulative distribution:

F(t)=P(T≤t)F(t)=P(T\le t)
=P ⁣(ZU/n≤t)=P\!\left(\frac{Z}{\sqrt{U/n}}\le t\right)
=P ⁣(Z≤tUn)=P\!\left(Z\le t\sqrt{\frac Un}\right)

Using the inequality this way...

−∞≤z≤tun,0≤u≤∞-\infty\le z\le t\sqrt{\frac un},\qquad 0\le u\le\infty

(Because U follows χ²ₙ.) We can do the double integral this way.

F(t)=∫0∞∫−∞tu/ng(z,u) dz duF(t)=\int_0^\infty\int_{-\infty}^{t\sqrt{u/n}}g(z,u)\,dz\,du
=∫0∞∫−∞tu/n12πe−z2/2un/2−1e−u/22n/2Γ(n/2) dz du=\int_0^\infty\int_{-\infty}^{t\sqrt{u/n}}\frac1{\sqrt{2\pi}}e^{-z^2/2}\frac{u^{n/2-1}e^{-u/2}}{2^{n/2}\Gamma(n/2)}\,dz\,du
=12π 2n/2Γ(n/2)∫0∞∫−∞tu/ne−z2/2un/2−1e−u/2 dz du=\frac1{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty\int_{-\infty}^{t\sqrt{u/n}}e^{-z^2/2}u^{n/2-1}e^{-u/2}\,dz\,du
=12π 2n/2Γ(n/2)∫0∞(∫−∞tu/ne−z2/2 dz)un/2−1e−u/2 du=\frac1{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty{\color{red}\Bigl(}\int_{-\infty}^{t\sqrt{u/n}}e^{-z^2/2}\,dz{\color{red}\Bigr)}u^{n/2-1}e^{-u/2}\,du

Let us call e to the power −z²/2 “h(z).”

e−z2/2=h(z)e^{-z^2/2}=h(z)
Preserved original img_025.pngPreserved original Korean source img_025.png
Differentiating and simplifying the density
F(t)=12π 2n/2Γ(n/2)∫0∞[H ⁣(tun)−H(−∞)]un/2−1e−u/2 duF(t)=\frac1{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty\left[H\!\left(t\sqrt{\frac un}\right)-H(-\infty)\right]u^{n/2-1}e^{-u/2}\,du

And now differentiate both sides with respect to t.

ddtF(t)=12π 2n/2Γ(n/2)∫0∞ddt[H ⁣(tun)−H(−∞)]un/2−1e−u/2 du{\color{red}\frac d{dt}}F(t)=\frac1{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty{\color{red}\frac d{dt}}\left[H\!\left(t\sqrt{\frac un}\right)-H(-\infty)\right]u^{n/2-1}e^{-u/2}\,du
f(t)=12π 2n/2Γ(n/2)∫0∞h ⁣(tun)un un/2−1e−u/2 duf(t)=\frac1{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty h\!\left(t\sqrt{\frac un}\right)\sqrt{\frac un}\,u^{n/2-1}e^{-u/2}\,du
=12π 2n/2Γ(n/2)∫0∞e−12unt2un↙Move √n outside too! un/2−1e−u/2 du=\frac1{\sqrt{2\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty e^{-\frac12\frac un t^2}\sqrt{\frac un}\,u^{n/2-1}e^{-u/2}\,du
=12nπ 2n/2Γ(n/2)∫0∞u un/2−1e−12unt2e−u/2 du=\frac1{\sqrt{2n\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty\sqrt u\,u^{n/2-1}e^{-\frac12\frac un t^2}e^{-u/2}\,du
=12nπ 2n/2Γ(n/2)∫0∞u(n+1)/2−1e−u2(1+t2/n) du=\frac1{\sqrt{2n\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty u^{(n+1)/2-1}e^{-\frac u2(1+t^2/n)}\,du

If we let u(1+t²/n)=y,

u(1+t2n)=y,u=y1+t2/n,du=11+t2/n dyu\left(1+\frac{t^2}n\right)=y,\qquad u=\frac y{1+t^2/n},\qquad du=\frac1{1+t^2/n}\,dy
=12nπ 2n/2Γ(n/2)∫0∞[y1+t2/n⏟Move this outside too!↘](n+1)/2−1e−y/2⋅11+t2/n↙Do not miss that these two factors combine here. dy=\frac1{\sqrt{2n\pi}\,2^{n/2}\Gamma(n/2)}\int_0^\infty\left[\frac y{1+t^2/n}\right]^{(n+1)/2-1}e^{-y/2}\cdot\frac1{1+t^2/n}\,dy
=12nπ 2n/2Γ(n/2)⋅1(1+t2/n)(n+1)/2∫0∞y(n+1)/2−1e−y/2 dy=\frac1{\sqrt{2n\pi}\,2^{n/2}\Gamma(n/2)}\cdot\frac1{(1+t^2/n)^{(n+1)/2}}\int_0^\infty y^{(n+1)/2-1}e^{-y/2}\,dy
=12nπ 2n/2Γ(n/2)⋅1(1+t2/n)(n+1)/2∫0∞2(n+1)/2Γ((n+1)/2) y(n+1)/2−1e−y/22(n+1)/2Γ((n+1)/2) dy=\frac1{\sqrt{2n\pi}\,2^{n/2}\Gamma(n/2)}\cdot\frac1{(1+t^2/n)^{(n+1)/2}}\int_0^\infty\frac{2^{(n+1)/2}\Gamma((n+1)/2)\,y^{(n+1)/2-1}e^{-y/2}}{2^{(n+1)/2}\Gamma((n+1)/2)}\,dy
=2(n+1)/2Γ((n+1)/2)2nπ 2n/2Γ(n/2)⋅1(1+t2/n)(n+1)/2∫0∞y(n+1)/2−1e−y/22(n+1)/2Γ((n+1)/2) dy=\frac{2^{(n+1)/2}\Gamma((n+1)/2)}{\sqrt{2n\pi}\,2^{n/2}\Gamma(n/2)}\cdot\frac1{(1+t^2/n)^{(n+1)/2}}\int_0^\infty\frac{y^{(n+1)/2-1}e^{-y/2}}{2^{(n+1)/2}\Gamma((n+1)/2)}\,dy

The integral’s density is the chi-squared distribution with n+1 degrees of freedom. Integrating over 0 to ∞ gives:

The integral is 1.

=Γ((n+1)/2)nπΓ(n/2)⋅1(1+t2/n)(n+1)/2=\frac{\Gamma((n+1)/2)}{\sqrt{n\pi}\Gamma(n/2)}\cdot\frac1{(1+t^2/n)^{(n+1)/2}}

Done!

∴ f(t)=Γ((n+1)/2)nπΓ(n/2)⋅1(1+t2/n)(n+1)/2\therefore\ f(t)=\frac{\Gamma((n+1)/2)}{\sqrt{n\pi}\Gamma(n/2)}\cdot\frac1{(1+t^2/n)^{(n+1)/2}}
Preserved original img_026.pngPreserved original Korean source img_026.png

Now let us briefly go over the properties of the t-distribution and finish this post.

First, the t-distribution is symmetric about the origin. (It is very easy to see that it is an even function.)

Also, scholars usually say that using the t-distribution is appropriate when “the sample size is less than 30.”

(Well, apparently some say 100, and others say 10,

but we can just think of 30 as the number they use on average.)

Next, the mean and variance of the t-distribution:

Mean: E(T) = 0, n > 1. (This also says that a t-distribution with one degree of freedom has no mean.)

Variance: Var(T) = n/n−2, n > 2. (Again, t-distributions with one or two degrees of freedom have no variance!!!!)

Finally, let us briefly cover how to read a t-distribution table.

(Although we might not need to,,,)

The t-distribution table writes the (1−α) quantile as t alpha,n,

which means Upper-tail probability alpha.

For example, if a table says t0.05,9 equals1.833... what does that mean?!

Right-tail5percent sketch at1.833

Hmm..... shall we do one example and finish?

(This problem comes from Walpole, et al., Probability and Statistics for Engineers and Scientists.)

Ex. 8.11 A chemical engineer claims that the yield of a batch process is 500 g per liter of raw material.

To demonstrate this, he selects and tests 25 batches every month.

It is agreed that his claim will be regarded as reasonable if the t value calculated from the test results lies between −t₀.₀₅ and t₀.₀₅.

If the results for the 25 batches give a sample mean of 518 and a standard deviation of 40 g, what conclusion can be drawn?

Assume that the population follows an approximately normal distribution.

Calculating the t value from the test results gives

t=(518−500)/(40/√25)=2.25

With 24 degrees of freedom, t₀.₀₅ = 1.711.

Since the calculated t value is greater than t₀.₀₅, we can say that the actual yield is greater than 500 g.

In fact, I have seen the t-distribution used a lot in t-tests,

but we have not discussed testing at all yet, so solving a t-test problem here does not seem quite right.

We will talk about testing later and solve one then.

n = [1., 2., 5., 100000000000000]
for i in n:                   # i = n
    x = np.linspace(-5, 5, 100)
    y = sc.t(i).pdf(x)
    plt.plot(x, y, linewidth=2.0, label = 'n=%s' % i)
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.title('Student-t Distribution')
# plt.savefig('5.Student-t Distribution.jpeg')

Source widget: cs / ColorScripter.

Student t distributions for increasing degrees of freedom

As the number gets bigger and bigger and bigger, it converges to the Gaussian.

Shall we check that on a computer too!??!

n = 5.
x = np.linspace(-0.3, 0.3, 100)
for i in range(3):
    y1 = sc.t(n).pdf(x)
    plt.plot(x, y1, linewidth=2.0, label = 'n=%s' % n)
    n = n * 5
 
y2 = sc.norm(0, 1).pdf(x)
plt.plot(x, y2, linewidth=1.0, label = 'Gaussian')
 
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.ylim(0.37, 0.4)
plt.title('Student-t & Gaussian Distribution')
# plt.savefig('6.Student-t & Gaussian Distribution.jpeg')

Source widget: cs / ColorScripter.

Student t and Gaussian distributions

After plotting it a few times, the middle seemed to catch up really, really fast,

so I adjusted the range a little.

After limiting it to that range,

let us make the number bigger and bigger again.

n = 50.
x = np.linspace(-0.3, 0.3, 100)
for i in range(3):
    y1 = sc.t(n).pdf(x)
    plt.plot(x, y1, linewidth=2.0, label = 'n=%s' % n)
    n = n * 5
 
y2 = sc.norm(0, 1).pdf(x)
plt.plot(x, y2, linewidth=1.0, label = 'Gaussian')
 
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.ylim(0.37, 0.4)
plt.title('Student-t & Gaussian Distribution(2)')
# plt.savefig('7.Student-t & Gaussian Distribution(2).jpeg')

Source widget: cs / ColorScripter.

Student t and Gaussian distributions second comparison

When I made n extremely large,

the curves looked completely overlapping to the eye,

so I chose a suitably large n at a sensible point......

Comments

Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.