Derivation of the Student's t-DistributionPhysicsBasic Statistics I Studied6Where the 'Student' name came from, why we ditch the Z-stat when σ is unknown, and a full derivation of the t-distribution PDF — plus properties and a worked example.
Where the 'Student' name came from, why we ditch the Z-stat when σ is unknown, and a full derivation of the t-distribution PDF — plus properties and a worked example.
·8 min read·gdpark
English translation of the original Statistics #6 post. Source notation and the author’s original claims are preserved. Separate explanatory notes identify and correct source errors.
First, why is this distribution called the Student t-distribution?
W. S. Grosset, who worked at a brewery, first discovered and used it while working there.
Apparently, he did not want other breweries to use it in the same way, so he published it under the pen name “student t.”
First, let us look at what the t-distribution is used for.
To estimate the population mean μ, we use the sample mean . The distribution related to is
and we use .
When we do not know , we cannot use the Z statistic,
so, as an alternative, we use , the T statistic.
The distribution followed by this T statistic is called the “t-distribution.”
The probability density function of this t-distribution is
, apparently.....
We cannot just skip past this without deriving it, can we????????
Actually, the derivation is the whole post, lol lol lol lol hehe.
Let us get into the derivation!
Rewriting T and introducing Z and U
First, let us change the form of T a little.
Phew...
After rewriting it this way, the numerator is
and we saw the denominator’s nS²/σ² when discussing the chi-squared distribution!
There, we proved that (n−1)S/σ² follows χ² with n−1 degrees of freedom, so we can say that nS²/σ² follows χ² with n degrees of freedom.
Then, for that T, substitute
Preserved original img_024.pngJoint density and cumulative distribution
As we have always done, let us calculate the cumulative distribution of this guy:
Since Z and U are statistically independent,
their joint probability density function (joint probability density function) is
If “statistically independent” or “joint probability density function” is unfamiliar, check the mathematical definition and have a think... It is not difficult.
Now, to calculate T’s cumulative distribution:
Using the inequality this way...
(Because U follows χ²ₙ.) We can do the double integral this way.
Let us call e to the power −z²/2 “h(z).”
Preserved original img_025.pngDifferentiating and simplifying the density
And now differentiate both sides with respect to t.
If we let u(1+t²/n)=y,
The integral’s density is the chi-squared distribution with n+1 degrees of freedom. Integrating over 0 to ∞ gives:
The integral is 1.
Done!
Preserved original img_026.png
Now let us briefly go over the properties of the t-distribution and finish this post.
First, the t-distribution is symmetric about the origin. (It is very easy to see that it is an even function.)
Also, scholars usually say that using the t-distribution is appropriate when “the sample size is less than 30.”
(Well, apparently some say 100, and others say 10,
but we can just think of 30 as the number they use on average.)
Next, the mean and variance of the t-distribution:
Mean: E(T) = 0, n > 1. (This also says that a t-distribution with one degree of freedom has no mean.)
Variance: Var(T) = n/n−2, n > 2. (Again, t-distributions with one or two degrees of freedom have no variance!!!!)
Finally, let us briefly cover how to read a t-distribution table.
(Although we might not need to,,,)
The t-distribution table writes the (1−α) quantile as ,
which means .
For example, if a table says ... what does that mean?!
Hmm..... shall we do one example and finish?
(This problem comes from Walpole, et al., Probability and Statistics for Engineers and Scientists.)
Ex. 8.11 A chemical engineer claims that the yield of a batch process is 500 g per liter of raw material.
To demonstrate this, he selects and tests 25 batches every month.
It is agreed that his claim will be regarded as reasonable if the t value calculated from the test results lies between −t₀.₀₅ and t₀.₀₅.
If the results for the 25 batches give a sample mean of 518 and a standard deviation of 40 g, what conclusion can be drawn?
Assume that the population follows an approximately normal distribution.
Calculating the t value from the test results gives
With 24 degrees of freedom, t₀.₀₅ = 1.711.
Since the calculated t value is greater than t₀.₀₅, we can say that the actual yield is greater than 500 g.
In fact, I have seen the t-distribution used a lot in t-tests,
but we have not discussed testing at all yet, so solving a t-test problem here does not seem quite right.
We will talk about testing later and solve one then.
n = [1., 2., 5., 100000000000000]
for i in n: # i = n
x = np.linspace(-5, 5, 100)
y = sc.t(i).pdf(x)
plt.plot(x, y, linewidth=2.0, label = 'n=%s' % i)
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.title('Student-t Distribution')
# plt.savefig('5.Student-t Distribution.jpeg')
Comments
Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.