English translation of the original Statistics #7 post . Original notation and claims are preserved. Separately labeled scientific notes qualify source errors and assumptions.
This time, we are on the last distribution: the F-distribution.
I think the reason we study the F-distribution after the t-distribution is
that although the statistic T follows a t-distribution, the square of T follows an F-distribution.
Of course, that does not mean it is a distribution followed only by the square of T.
Let us go more general—let’s go!
First, we need to define the random variable F.
Let me just throw this out there first.
, where U follows a chi-square distribution with degrees of freedom,
and V follows a chi-square distribution with degrees of freedom.
Since we had , it makes sense that the square of T follows an F-distribution. I will mention this once more at the very end.
And, to throw out the probability density function of the F-distribution first:
Source correction. The introductory formula above omits the exponent (r₁+r₂)/2 on 1+(r₁/r₂)f in the denominator. The later handwritten derivation includes it. The corrected density isg ( f ) = Γ ( r 1 + r 2 2 ) ( r 1 r 2 ) r 1 2 f r 1 2 − 1 Γ ( r 1 2 ) Γ ( r 2 2 ) ( 1 + r 1 r 2 f ) r 1 + r 2 2 , 0 < f < ∞ g(f)=\frac{\Gamma(\frac{r_1+r_2}2)(\frac{r_1}{r_2})^{\frac{r_1}2}f^{\frac{r_1}2-1}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)(1+\frac{r_1}{r_2}f)^{\frac{r_1+r_2}2}},\quad 0<f<\infty
Of course, this post is about deriving that F-distribution.
Then first, let us construct the joint density function of U and V in .
Assumption for the derivation. U and V are independent chi-square variables, with r₁ and r₂ degrees of freedom, respectively. Independence is needed to multiply their marginal densities to obtain h(u,v). The original explicitly states it later, in the F-distribution definition.
Referring to the chi-square probability density function ,
let us call the joint density function of U and V h(u,v).
(The way we construct the joint density function is the same as in the previous derivation of the t-distribution.)
And then the cumulative probability distribution.....
I want to write F(~), but that could be confusing,
so here I will use a different notation for the cumulative probability distribution!!!
Cumulative probability Cumulative probability (statistic ≤ f)
= P [ U / r 1 V / r 2 ≤ f ] = p [ r 2 r 1 1 V U ≤ f ] = p [ U ≤ r 1 r 2 V f ] =\operatorname{P}\!\left[\frac{U/r_1}{V/r_2}\le f\right]=p\!\left[\frac{r_2}{r_1}\frac1V U\le f\right]=p\!\left[U\le\frac{r_1}{r_2}Vf\right] Untouched native source witness
Since ,
from here on, I am going by hand;;;;
Typing this is too much~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Handwritten derivation, first sheet ∫ 0 ∞ ∫ 0 r 1 r 2 v f h ( u , v ) d u d v \int_0^\infty\!\int_0^{\frac{r_1}{r_2}vf}h(u,v)\,du\,dv = ∫ 0 ∞ ∫ 0 r 1 r 2 v f u r 1 2 − 1 v r 2 2 − 1 e − u + v 2 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 d u d v =\int_0^\infty\!\int_0^{\frac{r_1}{r_2}vf}\frac{u^{\frac{r_1}2-1}v^{\frac{r_2}2-1}e^{-\frac{u+v}2}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\,du\,dv = 1 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ ( ∫ 0 r 1 r 2 v f u r 1 2 − 1 e − u 2 d u ) v r 2 2 − 1 e − v 2 d v =\frac1{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty\left(\int_0^{\frac{r_1}{r_2}vf}u^{\frac{r_1}2-1}e^{-\frac u2}\,du\right)v^{\frac{r_2}2-1}e^{-\frac v2}\,dv Let the integrand here be ℓ(u).
= 1 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ ( L ( r 1 r 2 v f ) − L ( 0 ) ) v r 2 2 − 1 e − v 2 d v =\frac1{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty\left(\textcolor{red}{L(\frac{r_1}{r_2}vf)-L(0)}\right)v^{\frac{r_2}2-1}e^{-\frac v2}\,dv That gets us this far. Differentiate this straight away and move to the probability density function f!
d d f [ 1 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ ( L ( r 1 r 2 v f ) − L ( 0 ) ) v r 2 2 − 1 e − v 2 d v ] \textcolor{red}{\frac d{df}}\left[\frac1{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty\left(\textcolor{red}{L(\frac{r_1}{r_2}vf)-L(0)}\right)v^{\frac{r_2}2-1}e^{-\frac v2}\,dv\right] = 1 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ ℓ ( r 1 r 2 v f ) r 1 r 2 v ⋅ v r 2 2 − 1 e − v 2 d v =\frac1{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty\textcolor{blue}{\ell(\frac{r_1}{r_2}vf)\frac{r_1}{r_2}v}\cdot v^{\frac{r_2}2-1}e^{-\frac v2}\,dv Earlier, the integrand ℓ(u) was this.
ℓ ( u ) = u r 1 2 − 1 e − u 2 \ell(u)=u^{\frac{r_1}2-1}e^{-\frac u2} = 1 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ ( r 1 r 2 v f ) r 1 2 − 1 e − 1 2 ( r 1 r 2 v f ) ⋅ r 1 r 2 v ⋅ v r 2 2 − 1 e − v 2 d v =\frac1{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty\textcolor{blue}{(\frac{r_1}{r_2}vf)^{\frac{r_1}2-1}e^{-\frac12(\frac{r_1}{r_2}vf)}}\cdot\frac{r_1}{r_2}v\cdot v^{\frac{r_2}2-1}e^{-\frac v2}\,dv Just changing the order....
= 1 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ ( r 1 r 2 ) r 1 2 ⋅ v r 1 2 f r 1 2 − 1 e − 1 2 ( r 1 r 2 v f ) ⋅ v r 2 2 − 1 e − v 2 d v =\frac1{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty(\frac{r_1}{r_2})^{\frac{r_1}2}\cdot v^{\frac{r_1}2}f^{\frac{r_1}2-1}e^{-\frac12(\frac{r_1}{r_2}vf)}\cdot v^{\frac{r_2}2-1}e^{-\frac v2}\,dv The red arrow under the power of r₁/r₂: now take this out to the front.
= ( r 1 r 2 ) r 1 2 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ f r 1 2 − 1 v r 1 + r 2 2 − 1 e − 1 2 ( r 1 r 2 v f ) − v 2 d v =\frac{(\frac{r_1}{r_2})^{\frac{r_1}2}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty f^{\frac{r_1}2-1}v^{\frac{r_1+r_2}2-1}e^{-\frac12(\frac{r_1}{r_2}vf)-\frac v2}\,dv = ( r 1 r 2 ) r 1 2 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ f r 1 2 − 1 v r 1 + r 2 2 − 1 e − v 2 ( r 1 r 2 f + 1 ) d v =\frac{(\frac{r_1}{r_2})^{\frac{r_1}2}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty f^{\frac{r_1}2-1}v^{\frac{r_1+r_2}2-1}e^{-\frac v2(\frac{r_1}{r_2}f+1)}\,dv The green box marks the exponent for the substitution.
y = v 2 ( r 1 r 2 f + 1 ) ⇒ { d y = 1 2 ( r 1 r 2 f + 1 ) d v v = 2 y r 1 r 2 f + 1 d v = 2 r 1 r 2 f + 1 d y y=\frac v2(\frac{r_1}{r_2}f+1)\Rightarrow\begin{cases}dy=\frac12(\frac{r_1}{r_2}f+1)\,dv\\v=\frac{2y}{\frac{r_1}{r_2}f+1}\\dv=\frac2{\frac{r_1}{r_2}f+1}\,dy\end{cases} Untouched native source witness
Handwritten derivation, continuation = ( r 1 r 2 ) r 1 2 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ f r 1 2 − 1 [ 2 y r 1 r 2 f + 1 ] r 1 + r 2 2 − 1 e − y 2 r 1 r 2 f + 1 d y =\frac{(\frac{r_1}{r_2})^{\frac{r_1}2}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty f^{\frac{r_1}2-1}\left[\textcolor{green}{\frac{2y}{\frac{r_1}{r_2}f+1}}\right]^{\frac{r_1+r_2}2-1}e^{-y}\textcolor{green}{\frac2{\frac{r_1}{r_2}f+1}}\,dy Too many parentheses... let us expand them.
= ( r 1 r 2 ) r 1 2 Γ ( r 1 2 ) Γ ( r 2 2 ) 2 r 1 + r 2 2 ∫ 0 ∞ f r 1 2 − 1 2 r 1 + r 2 2 2 − 1 y r 1 + r 2 2 − 1 ( r 1 r 2 f + 1 ) − r 1 + r 2 2 + 1 e − y ⋅ 2 ⋅ ( r 1 r 2 f + 1 ) − 1 d y =\frac{(\frac{r_1}{r_2})^{\frac{r_1}2}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)2^{\frac{r_1+r_2}2}}\int_0^\infty f^{\frac{r_1}2-1}2^{\frac{r_1+r_2}2}2^{-1}y^{\frac{r_1+r_2}2-1}(\frac{r_1}{r_2}f+1)^{-\frac{r_1+r_2}2+1}e^{-y}\cdot2\cdot(\frac{r_1}{r_2}f+1)^{-1}\,dy The pink arrows take the power of f out to the front.
The red links: the powers of 2 cancel.
The blue links: the factors 2⁻¹ and 2 cancel.
The green links: the +1 and −1 powers cancel.
= ( r 1 r 2 ) r 1 2 f r 1 2 − 1 ( r 1 r 2 f + 1 ) − r 1 + r 2 2 Γ ( r 1 2 ) Γ ( r 2 2 ) ∫ 0 ∞ y r 1 + r 2 2 − 1 e − y d y =\frac{(\frac{r_1}{r_2})^{\frac{r_1}2}f^{\frac{r_1}2-1}(\frac{r_1}{r_2}f+1)^{-\frac{r_1+r_2}2}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)}\int_0^\infty y^{\frac{r_1+r_2}2-1}e^{-y}\,dy The Gamma function is defined as follows.
Γ ( n ) = ∫ 0 ∞ x n − 1 e − x d x \Gamma(n)=\int_0^\infty x^{n-1}e^{-x}\,dx That is, the red wavy underlined integral is Γ((r₁+r₂)/2).
= Γ ( r 1 + r 2 2 ) ( r 1 r 2 ) r 1 2 f r 1 2 − 1 Γ ( r 1 2 ) Γ ( r 2 2 ) ( r 1 r 2 f + 1 ) r 1 + r 2 2 =\frac{\Gamma(\frac{r_1+r_2}2)(\frac{r_1}{r_2})^{\frac{r_1}2}f^{\frac{r_1}2-1}}{\Gamma(\frac{r_1}2)\Gamma(\frac{r_2}2)(\frac{r_1}{r_2}f+1)^{\frac{r_1+r_2}2}} Untouched native source witness
That finishes the derivation.
Let us just pick up a few basic properties of the F-distribution and wrap this up!!!!
F-distribution definition U ∼ χ m 2 , V ∼ χ n 2 U\sim\chi_m^2,\quad V\sim\chi_n^2 When U and V are independent,
F = U / m V / n ∼ F ( m , n ) F=\frac{U/m}{V/n}\sim F(m,n) we say F follows an F-distribution with degrees of freedom (m,n). Just write ~ F(m,n).
Untouched native source witness
Likewise, this distribution has an F-distribution table.
Let us get a little practice reading it.
The F-distribution table denotes the (1−α)-quantile
by ,
which apparently means this:
Let us try an example with α = 0.05~
If we look up , it says 5.19.
Reading the table. α denotes the upper-tail probability. The value 5.19 is rounded: F₀.₀₅,₄,₅ ≈ 5.19217. The sketch shades the right tail, whose probability is 5%.
Probability 5% Untouched native source witness
Another, another, another, another, another property of the F-distribution:
if we switch the denominator and numerator in the definition of the F-distribution, we get an F-distribution with the degrees of freedom switched.
Reciprocal distribution F ∼ F m , n ⇒ 1 F ∼ F n , m F\sim F_{m,n}\quad\Rightarrow\quad\frac1F\sim F_{n,m} Let us also remember that this is what happens!
Untouched native source witness
Reciprocal quantile Taking the reciprocal of Fα,n,m gives
1 F α , n , m = F 1 − α , m , n \frac1{F_{\alpha,n,m}}=F_{1-\alpha,m,n} Untouched native source witness
So, for example, if we want to know , even without a table for α = 0.95,
we can find it like this as long as we have a table for α = 0.05!
Source reciprocal table calculation (source equality signs retained) F 0.95 , 4 , 5 = 1 F 0.05 , 5 , 4 = 1 read from the table = 1 6.26 = 0.16 F_{0.95,4,5}=\frac1{F_{0.05,5,4}}=\frac1{\text{read from the table}}=\frac1{6.26}=0.16 Untouched native source witness
Rounding note. The reciprocal quantile identity is exact, but the table values in the original calculation are rounded:F 0.95 , 4 , 5 = 1 F 0.05 , 5 , 4 ≈ 1 6.26 ≈ 0.16 F_{0.95,4,5}=\frac1{F_{0.05,5,4}}\approx\frac1{6.26}\approx0.16 The more precise lower cutoff is about 0.159845. The following sketch shades the right tail, whose probability is 95%.
Then this would be a picture like this.
To organize a little more about the F-distribution:
Two normal random samples X 1 , X 2 , ⋯ , X m from N ( μ 1 , σ 1 2 ) X_1,X_2,\cdots,X_m\quad\text{from }N(\mu_1,\sigma_1^2) are a random sample of size m, and
Y 1 , Y 2 , ⋯ , Y n from N ( μ 2 , σ 2 2 ) Y_1,Y_2,\cdots,Y_n\quad\text{from }N(\mu_2,\sigma_2^2) are a random sample of size n.
Untouched native source witness
The two random samples must also be independent!!!!
the result is .
Sampling assumptions. Within each sample, the observations are independent and identically distributed normal variables. The two samples are independent of each other. S₁² and S₂² are the usual unbiased sample variances, using divisors m−1 and n−1. Thus the numerator and denominator have m−1 and n−1 chi-square degrees of freedom. The normalized variance ratio shown here does not require equal population means or variances.
Why the sample-variance degrees are m−1 and n−1 F = S 1 2 σ 1 2 ⋅ m − 1 m − 1 S 2 2 σ 2 2 ⋅ n − 1 n − 1 F=\frac{\frac{S_1^2}{\sigma_1^2}\cdot\textcolor{red}{\frac{m-1}{m-1}}}{\frac{S_2^2}{\sigma_2^2}\cdot\textcolor{blue}{\frac{n-1}{n-1}}} If we transform the expression this way,
( m − 1 ) S 1 2 σ 1 2 ∼ χ ( m − 1 ) 2 and ( n − 1 ) S 2 2 σ 2 2 ∼ χ ( n − 1 ) 2 \frac{(m-1)S_1^2}{\sigma_1^2}\sim\chi_{(m-1)}^2\quad\text{and}\quad\frac{(n-1)S_2^2}{\sigma_2^2}\sim\chi_{(n-1)}^2 we covered these results earlier, so,
( m − 1 ) S 1 2 σ 1 2 = M , ( n − 1 ) S 2 2 σ 2 2 = N \frac{(m-1)S_1^2}{\sigma_1^2}=M,\quad\frac{(n-1)S_2^2}{\sigma_2^2}=N if we call these M and N,
F = M / ( m − 1 ) N / ( n − 1 ) ∼ F ( m − 1 , n − 1 ) F=\frac{M/(m-1)}{N/(n-1)}\sim F_{(m-1,n-1)} then that is indeed what F follows.
Untouched native source witness
And way~~~~~~back at the beginning,
for a statistic T that follows a t-distribution,
I said that the square of T follows an F-distribution.
Let us check the degrees of freedom too (although, of course, everyone has already done all~~~of this in their heads.... T.T).
The square of a t statistic T = Z U / n ∼ t ( n ) T=\frac Z{\sqrt{U/n}}\sim t_{(n)} That was what we had. Then
T 2 = Z 2 U / n T^2=\frac{Z^2}{U/n} is what this says,
Z 2 ∼ χ ( 1 ) 2 Z^2\sim\chi_{(1)}^2 which is almost common knowledge by now. Then,
T 2 = Z 2 U / n = Z 2 / 1 U / n ∼ F ( 1 , n ) T^2=\frac{Z^2}{U/n}=\frac{Z^2/1}{U/n}\sim F_{(1,n)} written this way, it is clear at a glance!
Untouched native source witness
Assumptions for the t-square relation. Z is standard normal and independent of U, where U follows a chi-square distribution with n degrees of freedom. Then T has n t degrees of freedom, and T² follows F(1,n).
That was just a summary of the F-distribution.
Now let us talk a little about testing..
r1 = [1, 2, 5, 10, 100]
r2 = [1, 1, 2, 1, 100]
r = list(zip(r1, r2))
for i, j in r: # i = n
x = np.linspace(0, 5, 1000)
y = sc.f(i, j).pdf(x)
plt.plot(x, y, linewidth=2.0, label = r'$r_1$=%s $r_2$=%s' % (i, j))
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.ylim(0, 2.5)
plt.title('F Distribution')
plt.savefig('8.F Distribution.jpeg')
Colored by Color Scripter cs
Comments
Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.