Derivation of the Gamma Distribution

Deriving the Gamma distribution from scratch by connecting it to the Poisson process — turns out it's just waiting for the α-th event instead of the first!!!!

Now it is time to move on to the gamma distribution,

and the Γ distribution is actually related to the exponential distribution!!!!

 

Why? Remember that the exponential distribution described the waiting time until the first Poisson event occurs in a Poisson process with mean λ

—that was what the exponential distribution was about, right???????????

 

 

 

Poisson process: 

Events occur randomly over time, and

1. The numbers of events occurring in nonoverlapping intervals are mutually independent;

2. The probability of one event occurring in a short interval is proportional to the length of that interval;

3. The probability of two or more events occurring in a short interval can be neglected;

4. If these conditions hold identically in every part of the time span,

we call it a Poisson process.

 

 

 

The gamma distribution

is the distribution representing the waiting time until α Poisson events have occurred!!!!

 

 

Let me throw out the definition first:

 

 That is the formula....

Original formula or English plot

 

Let us derive it briefly!!!!

We will follow the same approach as in our derivation of the exponential distribution!!

Let us think about the cumulative probability distribution.

 

 

 

The Poisson distribution

 was this.

Poisson distribution =

\[\frac{\lambda^xe^{-\lambda}}{x!}\]

 

 

This is a “Poisson distribution with an average of λ events per unit time,”

so a “Poisson distribution with an average of λt events over time t”

 looks like this.

Original formula or English plot

In this distribution, the probabilities of 0 events, 1 event, ~~~~~~~~~, α−1 events, and α events

 

will be g(0), g(1), ~~~~~~~~~~, g(α−1), and g(α). 

 

 

<Wait!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!>

Let us pause to point something out here.

In the gamma distribution, then, the things that seem meaningful are:

λ matters,

and α matters too!!!!

 

It matters which Poisson distribution we started with—how many events, on average, it represents over time t—

and it also matters how many events we are waiting for, does it not?!?!!?

 

I will reveal what these guys are at the very end.

 

 

 

 

Now, to change this into “the time until the α-th occurrence,”

I will add all the probabilities up to the α−1-th event (using the same idea as in the exponential-distribution derivation).

 

 

\[P(\text{fewer than }\alpha\text{ events occur})=g(0)+g(1)+\cdots+g(\alpha-1)\]

 

This means

 

 

\[P(\text{at least }\alpha\text{ events occur})=1-P(\text{fewer than }\alpha\text{ events occur})\]
\[=1-[g(0)+g(1)+\cdots+g(\alpha-1)]\]
\[=1-\sum_{x=0}^{\alpha-1}\frac{(\lambda t)^xe^{-\lambda t}}{x!}\]

 

and this means!!!!!!

Now, if we view the variable as time!!?

 

it becomes the “cumulative probability distribution of the waiting time for the α-th event”!

 

The probability of fewer than α events occurring

is the probability of t=1, the probability of t=2, the probability of t=3, the probability of t=4 ~~ and so on.

Does that make sense????

 

Then let us write it as the cumulative probability F(t)!

 

 

Original formula or English plot

Differentiating this will give us the probability density function!!!?!??!

Let us get differentiating!

 

Well... I will switch to photos.......

Hahahahahahaha, sorry, hahahahahahahahaha.

The equation editor is just too much...... sob, sob.

 

 

\[f(t)=\frac{d}{dt}F(t)=\frac{d}{dt}\left[1-\sum_{x=0}^{\alpha-1}\frac{(\lambda t)^x\cdot e^{-\lambda t}}{x!}\right]\mathord{.}\]
\[=-\frac{d}{dt}\sum_{x=0}^{\alpha-1}\frac{\lambda^x\cdot t^x\cdot e^{-\lambda t}}{x!}\]
\[=-\sum\frac{\lambda^x}{x!}\cdot\left(xt^{x-1}\cdot e^{-\lambda t}+t^x\cdot(-\lambda e^{-\lambda t})\right)\mathord{.}\]
\[=\sum\frac{\lambda^x}{x!}\cdot\left(\lambda t^x\cdot e^{-\lambda t}-xt^{x-1}\cdot e^{-\lambda t}\right)\mathord{.}\]
\[=\sum\left(\frac{\lambda\lambda^xt^xe^{-\lambda t}}{x!}-\frac{x\lambda\lambda^{x-1}t^{x-1}e^{-\lambda t}}{x!}\right)\mathord{.}\]
\[=\lambda e^{-\lambda t}\cdot\sum_{0}^{\alpha-1}\cdot\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)\mathord{.}\]
\[=\lambda e^{-\lambda t}\left(1+\sum_{x=1}^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)\right)\mathord{.}\]

The term left over when x=0 is taken out.

\[=\lambda e^{-\lambda t}+\lambda e^{-\lambda t}\cdot\underbrace{\sum_{1}^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)}\mathpunct{.}\]

Calculate this part separately.

\[\sum_{1}^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)=\sum_{1}^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{(\lambda t)^{x-1}}{(x-1)!}\right)\mathpunct{.}\]
\[=\bigl(\frac{(\lambda t)^1}{1!}\]
\[-\frac{(\lambda t)^0}{0!}\bigr)\]
\[+\bigl(\frac{(\lambda t)^2}{2!}\]
\[-\frac{(\lambda t)^1}{1!}\bigr)\mathpunct{.}\]
\[+\bigl(\frac{(\lambda t)^3}{3!}\]
\[-\frac{(\lambda t)^2}{2!}\bigr)\]
\[+\cdots+\bigl(\frac{(\lambda t)^{\alpha-2}}{(\alpha-2)!}\]
\[-\frac{(\lambda t)^{\alpha-3}}{(\alpha-3)!}\bigr)\]
\[+\bigl(\frac{(\lambda t)^{\alpha-1}}{(\alpha-1)!}\]
\[-\text{[right-edge clipped strokes]}\bigr)\]

Most of them cancel... only two terms survive.

The complete visible bottom-right RHS, including its equals sign, fraction, denominator parentheses and factorial, and minus one, is transcribed below. The left-side sum and edge connecting strokes remain clipped and unknown; no complete bottom equation is reconstructed.

\[=\frac{(\lambda t)^{\alpha-1}}{(\alpha-1)!}-1\]

Source limitation: The original clips strokes at its right and bottom edges; only the visible text is transcribed.

 

 

 

Original formula or English plot

 

 That is how far we got!!

Original formula or English plot

 

Writing that factorial in the denominator this way makes the expression defined only when (α−1) is zero or a natural number,

so let us use the gamma function to write it more generally!!!

(The derivation of the formula is here: http://gdpresent.blog.me/220581881465 !!!!)

 

 

 

Here I will just throw out the definition of the gamma function!

 

 

Original formula or English plot

That is it. So now we can express the formula above using the gamma function!

 

 

 

\[f(t)=\frac{\lambda(\lambda t)^{\alpha-1}e^{-\lambda t}}{(\alpha-1)!}\]
\[=\frac{\lambda(\lambda t)^{\alpha-1}e^{-\lambda t}}{\Gamma(\alpha)}\]
\[=\frac{\lambda^\alpha t^{\alpha-1}e^{-\lambda t}}{\Gamma(\alpha)}\]

← We have now derived the definition thrown out at the start of this post.

\[=\frac{t^{\alpha-1}e^{-\lambda t}}{\lambda^{-\alpha}\cdot\Gamma(\alpha)}\]

← It can also be expressed this way.

\[\frac1\lambda=\theta\]

If we express it this way,

\[f(t)=\frac{t^{\alpha-1}e^{-t/\theta}}{\theta^\alpha\Gamma(\alpha)}\]

← It can also be expressed this way.

 

 

Let us point this out once more?!?!?! What does f(t) mean?!?!?!?!?

 

It will be “the probability that the time until the α-th event occurs is t!!!!!”!!~

 

cf.) Of course, if we put α=1 in f(t), it becomes “the probability that the waiting time until the first event is t,”

    or, put another way, “the probability that the first event occurs at time t.”

    In other words, it reduces back to the exponential distribution.... sob!

 

Now, since we are at the end, I will reveal what alpha and gamma are.

These are the two “parameters” of the gamma distribution. 

 

Parameters~??~??~?

What-a-meters~?!?

 

Sorry, sorry.

How should I explain parameters???

Calling them “characteristics of the population” probably does not make it click, does it?

 

The Gaussian distribution that we know so very well also has two parameters.

They are the mean and variance (the μ and σ values in the Gaussian function determine that particular Gaussian’s shape and characteristics).

 

Likewise, the parameters determining the shape and characteristics of a particular gamma function here are lambda and alpha!!!

So what I am trying to say is:

 

 following a gamma distribution like this

Original formula or English plot

means we say “the random variable t follows a gamma distribution with parameters (α, λ)”~~~

Let us move on~~~ (I will skip shape parameter and scale parameter... go study them yourselves!)

 

 

First, let us check whether the sum over the whole domain of the Gamma Distribution is 1.

The sum over the entire range of a probability distribution must be 1, right?!~

Because all probabilities must add up to 1! Let us go, go, go!

 

 

\[f(t)=\frac{t^{\alpha-1}\cdot e^{-\lambda t}\mathpunct{\raisebox{0.6em}{.}}}{\lambda^{-\alpha}\Gamma(\alpha)}\]

Since t≥0, the integration interval is from 0 to ∞.

\[\int_0^\infty f(t)\,dt=\int_0^\infty\frac{t^{\alpha-1}e^{-\lambda t}}{\lambda^{-\alpha}\Gamma(\alpha)}\cdot dt\mathord{.}\]
\[\lambda t=y\quad\Rightarrow\quad dt=\frac1\lambda\,dy\mathpunct{.}\]
\[=\int_0^\infty\cdot\frac{t^{\alpha-1}\cdot e^{-y}}{\lambda^{-\alpha}\Gamma(\alpha)}\cdot\frac1\lambda\,dy\mathord{.}\]
\[=\int_0^\infty\cdot\frac{t^{\alpha-1}e^{-y}}{\lambda^{1-\alpha}\cdot\Gamma(\alpha)}\,dy\]
\[=\int_0^\infty\frac{(\lambda t)^{\alpha-1}\cdot e^{-y}}{\Gamma(\alpha)}\cdot dy\]
\[=\frac{\int_0^\infty\cdot y^{\alpha-1}e^{-y}\,dy\mathpunct{.}}{\Gamma(\alpha)}\]

← The gamma-function definition appears again!

\[=\frac{\Gamma(\alpha)}{\Gamma(\alpha)}=1\]

Check finished~

 

 

 

Let us calculate μ and σ² again with the moment generating function.

When we first wrote the probability distribution, we used t (time) as its variable. But in the MGF, t is already the MGF variable! So let us change the t we have used so far into x!

\[f(x)=\frac{x^{\alpha-1}e^{-\lambda x}\mathpunct{\raisebox{0.6em}{.}}}{\lambda^{-\alpha}\Gamma(\alpha)}\]

(It was t before.)

To bring out the moment generating function:

\[M(t)=\int_0^\infty e^{tx}\cdot f(x)\,dx\]
\[=\int_0^\infty e^{tx}\cdot\frac{x^{\alpha-1}e^{-\lambda x}}{\lambda^{-\alpha}\Gamma(\alpha)}\,dx\mathord{.}\]
\[=\int_0^\infty\frac{x^{\alpha-1}e^{-(\lambda-t)x}}{\lambda^{-\alpha}\Gamma(\alpha)}\,dx\mathord{.}\]
\[\textcolor{red}{\left\{\begin{gathered}(\lambda-t)x=y\\dx=\frac1{\lambda-t}\,dy\\x^{\alpha-1}=\frac{y^{\alpha-1}}{(\lambda-t)^{\alpha-1}}\end{gathered}\right.}\]
\[=\int_0^\infty\frac{\frac{y^{\alpha-1}}{(\lambda-t)^{\alpha-1}}\cdot e^{-y}}{\lambda^{-\alpha}\cdot\Gamma(\alpha)}\cdot\frac1{\lambda-t}\,dy\]
\[=\int_0^\infty\frac{y^{\alpha-1}\cdot e^{-y}}{(\lambda-t)^\alpha\cdot\lambda^{-\alpha}\cdot\Gamma(\alpha)}\cdot dy\]

The gamma-function definition appears again!

\[=\frac{\cancel{\Gamma(\alpha)}}{(\lambda-t)^\alpha\lambda^{-\alpha}\cdot\cancel{\Gamma(\alpha)}}=\frac1{\left(\frac{\lambda-t}{\lambda}\right)^\alpha}=\frac1{(1-t/\lambda)^\alpha}=\left(1-\frac t\lambda\right)^{-\alpha}\mathpunct{\raisebox{0.6em}{.}}\]

 

 

 

 

 

\[M(t)=\left(1-\frac t\lambda\right)^{-\alpha}\]
\[\frac d{dt}\cdot M(t)=M'(t)=-\alpha\cdot\left(1-\frac t\lambda\right)^{-\alpha-1}\cdot\left(-\frac1\lambda\right)=\frac\alpha\lambda\left(1-\frac t\lambda\right)^{-\alpha-1}\]
\[M'(0)=\frac\alpha\lambda(1)=\textcolor{#873c50}{\underline{\textcolor{black}{\frac\alpha\lambda\quad\leftarrow\mu\mathord{.}}}}\]

OK~

\[\frac d{dt}\cdot M'(t)=\frac d{dt}\left(\frac\alpha\lambda\left(1-\frac t\lambda\right)^{-\alpha-1}\right)=\frac\alpha\lambda(-\alpha-1)\cdot\left(1-\frac t\lambda\right)^{-\alpha-2}\cdot\left(-\frac1\lambda\right)\]
\[=\frac\alpha{\lambda^2}(\alpha+1)\left(1-\frac t\lambda\right)^{-\alpha-2}\mathpunct{\raisebox{0.6em}{.}}\]
\[M''(0)=\frac\alpha{\lambda^2}(\alpha+1)=\frac{\alpha(\alpha+1)}{\lambda^2}=E(x^2)\mathord{.}\]
\[\therefore\sigma^2=E(x^2)-E(x)^2=\frac{\alpha(\alpha+1)}{\lambda^2}-\left(\frac\alpha\lambda\right)^2\]
\[=\frac{\alpha^2+\alpha-\alpha^2}{\lambda^2}=\textcolor{#873c50}{\underline{\textcolor{black}{\frac\alpha{\lambda^2}\quad\leftarrow\sigma^2\mathpunct{\raisebox{0.6em}{.}}}}}\]

All right~

 

 

 

 

 

 

 

 

 

Finally, let us organize the “additivity” property before moving on!

If a random variable follows the displayed formula, we say “the random variable x follows a gamma distribution with parameters (α, λ),” remember????

Original formula or English plot

And we checked that it reduces to the exponential distribution when α=1.

 

 

Now suppose we have α random variables X₁, X₂, ..... , Xα, 

each mutually independent and following an exponential distribution with parameter λ!!!

(You can already get a feel for what I am going to say, right??? If you have a gamma feeling, it is the gamma distribution....!!! Sorry!!!!!!!!!!!!!!!!)

 

 

You have not forgotten, right??????

The exponential distribution with parameter λ  is the displayed formula!

Original formula or English plot

 

 

Oh, whatever... you have all got the idea already.....

Let me give the conclusion first,,,,

 

 

Original formula or English plot

Additivity means that the random variable Y defined this way follows a gamma distribution with parameters (α, λ).

I will prove it!

 

It is very simple for something called a proof, but I thought it would be good to know the principle!!!

 

 

 If we bring out the moment generating function of the random variable Y and keep going,

Original formula or English plot

 

From here on, bye-bye, equation editor~

I will rely on the iPad.

 

 

 

\[Y=X_1+X_2+X_3+\cdots+X_\alpha\]
\[M_Y(t)=E\left(e^{(X_1+X_2+\cdots+X_\alpha)t}\right)\]
\[=E\left(e^{X_1t}\cdot e^{X_2t}\cdots e^{X_\alpha t}\right)\]
\[=E(e^{X_1t})\cdot E(e^{X_2t})\cdots E(e^{X_\alpha t})\]

Because we assumed independence!

\[=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_\alpha}(t)\]

We said that X₁, X₂, ..., Xα each follow an exponential distribution with λ, and already checked that its moment generating function is λ/(λ−t).

\[\frac\lambda{\lambda-t}\]
\[=\left(\frac{\lambda-t}{\lambda}\right)^{-1}\]
\[=\left(1-\frac t\lambda\right)^{-1}\]
\[M_Y(t)=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_\alpha}(t)\]
\[=\left(1-\frac t\lambda\right)^{-1}\cdot\left(1-\frac t\lambda\right)^{-1}\cdots\left(1-\frac t\lambda\right)^{-1}\]
\[=\left(1-\frac t\lambda\right)^{-\alpha}\]

This proves the statement above.

Let us prove another very similar theorem.

Suppose the k random variables X₁, X₂, ..., Xk are mutually independent and each Xi follows a gamma distribution with (ri, λ).

\[Y=X_1+X_2+\cdots+X_k\]
\[M_Y(t)=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_k}(t)\]

The MGF of a gamma distribution with (α, λ) is

\[\left(1-\frac t\lambda\right)^{-\alpha}\]

Since Xi follows a gamma distribution with (ri, λ),

\[M_{X_i}(t)=\left(1-\frac t\lambda\right)^{-r_i}\]
\[M_Y(t)=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_k}(t)\]
\[=\left(1-\frac t\lambda\right)^{-r_1}\cdot\left(1-\frac t\lambda\right)^{-r_2}\cdots\left(1-\frac t\lambda\right)^{-r_k}\]
\[=\left(1-\frac t\lambda\right)^{-(r_1+r_2+\cdots+r_k)}\]

Therefore, for mutually independent Xi, each following a gamma distribution with (ri, λ), Y=X₁+X₂+···+Xk follows a gamma distribution with (r₁+r₂+···+rk, λ).

\[X\sim\mathrm{Gamma}(r,\lambda)\quad\Rightarrow\quad cX\sim\mathrm{Gamma}(cr,\lambda)\]

Even scalar multiplication can all be explained!

Correction to this original green claim: for c>0, cX has shape r and rate λ/c, so cX∼Gamma(r,λ/c). Gamma(cr,λ) instead describes a sum of c independent copies for positive integer c. This shared correction applies to both img_117 and img_123.

Source limitation: The historical scaling assertion is reproduced here; see the adjacent correction.

 

 

 

 

The end~~~~

 

 

 

 

 

 

↓ A slightly prettier version on the iPad! Haha.

 

 

 

 

 

Gamma Distribution

\[\text{Poisson distribution: }\frac{\lambda^xe^{-\lambda}}{x!}\]

The probability of x events occurring when the average number of events per unit time is λ.

Visible small right-margin fragment: “per unit time, on average λ events occur…”; unseen trailing ink is unknown.

\[\text{let }g(x)=\frac{(\lambda t)^xe^{-\lambda t}}{x!}\]

The probability of x events occurring over time t when the average number of events over time t is λt.

Visible small right-margin fragment: “over time t, on average λt events occur…”; unseen trailing ink is unknown.

Then, with an average of λt events over time t, the probability of 0 events occurring = g(0).

With an average of λt events over time t, the probability of 1 event occurring = g(1).

With an average of λt events over time t, the probability of 2 events occurring = g(2).

\[\vdots\]

With an average of λt events over time t, the probability of α−1 events occurring = g(α−1).

With an average of λt events over time t, the probability of α events occurring = g(α).

Therefore, with an average of λt events over time t, the probability of fewer than α events occurring is

\[g(0)+g(1)+\cdots+g(\alpha-1)\]

Or? With an average of λt events over time t, the probability of α or more events occurring is

\[1-[g(0)+g(1)+\cdots+g(\alpha-1)]\]
\[=1-\sum_{x=0}^{\alpha-1}g(x)\]
\[=1-\sum_{x=0}^{\alpha-1}\frac{(\lambda t)^xe^{-\lambda t}}{x!}\]

Interpretation: with an average of λt events over time t, the probability that the waiting time until the α-th event is smaller than t.

All right!

\[P(T\leq t)=1-\sum_{x=0}^{\alpha-1}\frac{(\lambda t)^xe^{-\lambda t}}{x!}\]
\[\text{Cumulative probability distribution }=1-\sum_{x=0}^{\alpha-1}\frac{(\lambda t)^xe^{-\lambda t}}{x!}\]

↓ Differentiate!

\[\text{Probability density function }=\frac d{dt}\left(1-\sum_{x=0}^{\alpha-1}\frac{(\lambda t)^xe^{-\lambda t}}{x!}\right)\]

On the next page...

Source limitation: Native small right-edge notes are clipped; no unseen text inferred.

\[\frac d{dt}\left(1-\sum_0^{\alpha-1}\frac{(\lambda t)^xe^{-\lambda t}}{x!}\right)=-\sum_0^{\alpha-1}\frac{\lambda^x}{x!}\cdot\frac d{dt}(t^x\cdot e^{-\lambda t})\]
\[=-\sum_0^{\alpha-1}\frac{\lambda^x}{x!}\left(xt^{x-1}\cdot e^{-\lambda t}-\lambda t^xe^{-\lambda t}\right)\]
\[=\sum_0^{\alpha-1}\frac{\lambda^x}{x!}\left(\lambda t^xe^{-\lambda t}-xt^{x-1}e^{-\lambda t}\right)\]
\[=\sum_0^{\alpha-1}\left(\frac{\lambda^x\cdot\lambda t^xe^{-\lambda t}}{x!}-\frac{x\lambda\cdot\lambda^{x-1}\cdot t^{x-1}e^{-\lambda t}}{x!}\right)\]
\[=\sum_0^{\alpha-1}\left(\frac{\lambda^x\cdot\textcolor{red}{\lambda}t^x\textcolor{red}{e^{-\lambda t}}}{x!}-\frac{x\textcolor{red}{\lambda}\cdot\lambda^{x-1}\cdot t^{x-1}\textcolor{red}{e^{-\lambda t}}}{x!}\right)\]
\[=\textcolor{red}{\lambda e^{-\lambda t}}\sum_0^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)\]
\[=\lambda e^{-\lambda t}\left(1+\sum_{\textcolor{blue}{1}}^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)\right)\]
\[=\lambda e^{-\lambda t}+\lambda e^{-\lambda t}\sum_1^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)\]
\[=\lambda e^{-\lambda t}+\lambda e^{-\lambda t}\left(\frac{(\lambda t)^{\alpha-1}}{(\alpha-1)!}-1\right)\]
\[=\frac{\lambda(\lambda t)^{\alpha-1}e^{-\lambda t}}{(\alpha-1)!}\]

Since Γ(n)=(n−1)!

\[=\frac{\lambda(\lambda t)^{\alpha-1}e^{-\lambda t}}{\Gamma(\alpha)}\]
\[\textcolor{green}{\sum_1^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{x(\lambda t)^{x-1}}{x!}\right)}\]
\[\textcolor{green}{=\sum_1^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{(\lambda t)^{x-1}}{(x-1)!}\right)}\]
\[\textcolor{green}{=\left(\textcolor{black}{\frac{(\lambda t)^1}{1!}}-\textcolor{green}{\frac{(\lambda t)^0}{0!}}\right)+\left(\textcolor{blue}{\frac{(\lambda t)^2}{2!}}-\textcolor{black}{\frac{(\lambda t)^1}{1!}}\right)+\left(\textcolor{green}{\frac{(\lambda t)^3}{3!}}-\textcolor{blue}{\frac{(\lambda t)^2}{2!}}\right)}\]
\[\textcolor{green}{+\cdots+\left(\textcolor{red}{\frac{(\lambda t)^{\alpha-2}}{(\alpha-2)!}}-\textcolor{green}{\frac{(\lambda t)^{\alpha-3}}{(\alpha-3)!}}\right)+\left(\textcolor{green}{\frac{(\lambda t)^{\alpha-1}}{(\alpha-1)!}}-\textcolor{red}{\frac{(\lambda t)^{\alpha-2}}{(\alpha-2)!}}\right)}\]
\[\textcolor{green}{=\left(\textcolor{black}{\cancel{\frac{(\lambda t)^1}{1!}}}-\textcolor{green}{\frac{(\lambda t)^0}{0!}}\right)+\left(\textcolor{blue}{\cancel{\frac{(\lambda t)^2}{2!}}}-\textcolor{black}{\cancel{\frac{(\lambda t)^1}{1!}}}\right)+\left(\textcolor{green}{\frac{(\lambda t)^3}{3!}}-\textcolor{blue}{\cancel{\frac{(\lambda t)^2}{2!}}}\right)}\]
\[\textcolor{green}{+\cdots+\left(\textcolor{red}{\frac{(\lambda t)^{\alpha-2}}{(\alpha-2)!}}-\textcolor{green}{\frac{(\lambda t)^{\alpha-3}}{(\alpha-3)!}}\right)+\left(\textcolor{green}{\frac{(\lambda t)^{\alpha-1}}{(\alpha-1)!}}-\textcolor{red}{\frac{(\lambda t)^{\alpha-2}}{(\alpha-2)!}}\right)}\]
\[\textcolor{green}{\therefore\sum_1^{\alpha-1}\left(\frac{(\lambda t)^x}{x!}-\frac{(\lambda t)^{x-1}}{(x-1)!}\right)=\frac{(\lambda t)^{\alpha-1}}{(\alpha-1)!}-1}\]
\[\textcolor{red}{\text{cf. Gamma-function definition: }\Gamma(n)=\int_0^\infty x^{n-1}e^{-x}\,dx=(n-1)!}\]
x=0 term taken out
\[=\frac{\lambda^\alpha t^{\alpha-1}e^{-\lambda t}}{\Gamma(\alpha)}\]
\[=\frac{t^{\alpha-1}e^{-\lambda t}}{\lambda^{-\alpha}\Gamma(\alpha)}=\frac{t^{\alpha-1}e^{-t/\theta}}{\theta^\alpha\Gamma(\alpha)}\quad\left(\theta=\frac1\lambda\right)\]

It is over!

\[\therefore P(t)=\frac{\lambda^\alpha t^{\alpha-1}e^{-\lambda t}}{\Gamma(\alpha)}\]

Meaning: with an average of λt events over time t, the probability that the waiting time until the α-th event is t.

Check 1. Does integrating the probability density function over its entire range give 1?

\[\int_0^\infty P(t)\,dt=\int_0^\infty\frac{t^{\alpha-1}e^{-\lambda t}}{\lambda^{-\alpha}\Gamma(\alpha)}\,dt\]
\[=\int\frac{\left(\frac1\lambda y\right)^{\alpha-1}e^{-y}}{\lambda^{-\alpha}\Gamma(\alpha)}\cdot\frac1\lambda\,dy\]
\[=\int\frac{\textcolor{red}{\cancel{\textcolor{black}{\lambda^{-\alpha+1}}}}y^{\alpha-1}e^{-y}}{\textcolor{red}{\cancel{\textcolor{black}{\lambda^{1-\alpha}}}}\Gamma(\alpha)}\,dy\]
\[=\frac{\int_0^\infty y^{\alpha-1}e^{-y}\,dy}{\Gamma(\alpha)}\]
\[=\frac{\Gamma(\alpha)}{\Gamma(\alpha)}\]
\[=1\]

Check finished!

\[\textcolor{blue}{\text{* }\lambda t=y\quad\Rightarrow\quad dt=\frac1\lambda\,dy}\]
\[\textcolor{red}{\begin{gathered}\text{cf. Gamma-function definition: }\\\Gamma(n)=\int_0^\infty x^{n-1}e^{-x}\,dx\\=(n-1)!\end{gathered}}\]

Check 2. μ and σ²

Wait! Since the moment generating function uses t as its variable, let us call the gamma-distribution variable x!

\[f(x)=\frac{x^{\alpha-1}e^{-\lambda x}}{\lambda^{-\alpha}\Gamma(\alpha)}\]
\[M(t)=\int_0^\infty e^{tx}f(x)\,dx\]
\[=\int_0^\infty e^{tx}\frac{x^{\alpha-1}e^{-\lambda x}}{\lambda^{-\alpha}\Gamma(\alpha)}\,dx\]
\[=\int_0^\infty\frac{x^{\alpha-1}e^{-(\lambda-t)x}}{\lambda^{-\alpha}\Gamma(\alpha)}\,dx\]
\[=\int_0^\infty\frac{\frac{y^{\alpha-1}}{(\lambda-t)^{\alpha-1}}\cdot e^{-y}}{\lambda^{-\alpha}\Gamma(\alpha)}\cdot\frac1{\lambda-t}\,dy\]
\[=\frac{\int_0^\infty y^{\alpha-1}e^{-y}\,dy}{(\lambda-t)^\alpha\lambda^{-\alpha}\Gamma(\alpha)}\]
\[=\frac{\textcolor{red}{\cancel{\textcolor{black}{\Gamma(\alpha)}}}}{(\lambda-t)^\alpha\lambda^{-\alpha}\textcolor{red}{\cancel{\textcolor{black}{\Gamma(\alpha)}}}}\]
\[=\frac1{(\lambda-t)^\alpha\lambda^{-\alpha}}\]
\[=\frac1{(1-t/\lambda)^\alpha}\]
\[M(t)=\left(1-\frac t\lambda\right)^{-\alpha}\]
\[M'(t)=-\alpha\left(1-\frac t\lambda\right)^{-\alpha-1}\left(-\frac1\lambda\right)\]
\[=\frac\alpha\lambda\left(1-\frac t\lambda\right)^{-(\alpha+1)}\]
\[\textcolor{green}{\mu=M'(0)=\frac\alpha\lambda}\]
\[M''(t)=\frac d{dt}\left(\frac\alpha\lambda\left(1-\frac t\lambda\right)^{-(\alpha+1)}\right)\]
\[=\frac\alpha\lambda(-\alpha-1)\left(1-\frac t\lambda\right)^{-(\alpha+2)}\cdot\left(-\frac1\lambda\right)\]
\[=\frac{\alpha(\alpha+1)}{\lambda^2}\cdot\left(1-\frac t\lambda\right)^{-(\alpha+2)}\]
\[\textcolor{green}{\sigma^2=M''(0)-M'(0)^2}\]
\[\textcolor{green}{=\frac{\alpha(\alpha+1)}{\lambda^2}-\left(\frac\alpha\lambda\right)^2}\]
\[\textcolor{green}{=\frac\alpha{\lambda^2}}\]
\[\textcolor{red}{\begin{gathered}(\lambda-t)x=y\\dx=\frac1{\lambda-t}\,dy\\x^{\alpha-1}=\frac{y^{\alpha-1}}{(\lambda-t)^{\alpha-1}}\end{gathered}}\]
\[\textcolor{red}{\begin{gathered}\text{cf. Gamma-function definition: }\\\Gamma(n)=\int_0^\infty x^{n-1}e^{-x}\,dx\\=(n-1)!\end{gathered}}\]
\[Y=X_1+X_2+X_3+\cdots+X_\alpha\]
\[M_Y(t)=E\left(e^{(X_1+X_2+\cdots+X_\alpha)t}\right)\]
\[=E\left(e^{X_1t}\cdot e^{X_2t}\cdots e^{X_\alpha t}\right)\]
\[=E(e^{X_1t})\cdot E(e^{X_2t})\cdots E(e^{X_\alpha t})\]

Because we assumed independence!

\[=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_\alpha}(t)\]

We said that X₁, X₂, ..., Xα each follow an exponential distribution with λ, and already checked that its moment generating function is λ/(λ−t).

\[\frac\lambda{\lambda-t}\]
\[=\left(\frac{\lambda-t}{\lambda}\right)^{-1}\]
\[=\left(1-\frac t\lambda\right)^{-1}\]
\[M_Y(t)=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_\alpha}(t)\]
\[=\left(1-\frac t\lambda\right)^{-1}\cdot\left(1-\frac t\lambda\right)^{-1}\cdots\left(1-\frac t\lambda\right)^{-1}\]
\[=\left(1-\frac t\lambda\right)^{-\alpha}\]

This proves the statement above.

Let us prove another very similar theorem.

Suppose the k random variables X₁, X₂, ..., Xk are mutually independent and each Xi follows a gamma distribution with (ri, λ).

\[Y=X_1+X_2+\cdots+X_k\]
\[M_Y(t)=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_k}(t)\]

The MGF of a gamma distribution with (α, λ) is

\[\left(1-\frac t\lambda\right)^{-\alpha}\]

Since Xi follows a gamma distribution with (ri, λ),

\[M_{X_i}(t)=\left(1-\frac t\lambda\right)^{-r_i}\]
\[M_Y(t)=M_{X_1}(t)\cdot M_{X_2}(t)\cdots M_{X_k}(t)\]
\[=\left(1-\frac t\lambda\right)^{-r_1}\cdot\left(1-\frac t\lambda\right)^{-r_2}\cdots\left(1-\frac t\lambda\right)^{-r_k}\]
\[=\left(1-\frac t\lambda\right)^{-(r_1+r_2+\cdots+r_k)}\]

Therefore, for mutually independent Xi, each following a gamma distribution with (ri, λ), Y=X₁+X₂+···+Xk follows a gamma distribution with (r₁+r₂+···+rk, λ).

\[X\sim\mathrm{Gamma}(r,\lambda)\quad\Rightarrow\quad cX\sim\mathrm{Gamma}(cr,\lambda)\]

Even scalar multiplication can all be explained!

Correction to this original green claim: for c>0, cX has shape r and rate λ/c, so cX∼Gamma(r,λ/c). Gamma(cr,λ) instead describes a sum of c independent copies for positive integer c. This shared correction applies to both img_117 and img_123.

Source limitation: The historical scaling assertion is reproduced here; see the adjacent correction.

 

 

 

 

 

 

 

 

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
import numpy as np
import matplotlib.pyplot as plt
import os
import scipy.stats as sc
 
 
pdp = []
alpha = 1.0
llambda = 0.5
scale= 1./llambda
x = np.linspace(0, 30, 100)
y = sc.gamma.pdf(x, a=alpha, scale=scale)
plt.plot(x, y, linewidth=2.0, label = r'$\alpha =$1   $\lambda =$0.5')
 
pdp = []
alpha = 2.0
llambda = 0.5
scale = 1./llambda
x = np.linspace(0, 30, 100)
y = sc.gamma.pdf(x, a=alpha, scale=scale)
plt.plot(x, y, linewidth=2.0, label=r'$\alpha =$2   $\lambda =$0.5')
 
pdp = []
alpha = 3.0
llambda = 0.5
loc = alpha / llambda
scale = 1./llambda
x = np.linspace(0, 30, 100)
y = sc.gamma.pdf(x, a=alpha, scale=scale)
plt.plot(x, y, linewidth=2.0, label=r'$\alpha =$3   $\lambda =$0.5')
 
pdp = []
alpha = 5.0
llambda = 1.0
scale = 1./llambda
x = np.linspace(0, 30, 100)
y = sc.gamma.pdf(x, a=alpha, scale=scale)
plt.plot(x, y, linewidth=2.0, label=r'$\alpha =$5   $\lambda =$1')
 
pdp = []
alpha = 9.0
llambda = 2.0
scale = 1./llambda
x = np.linspace(0, 30, 100)
y = sc.gamma.pdf(x, a=alpha, scale=scale)
plt.plot(x, y, linewidth=2.0, label=r'$\alpha =$9   $\lambda =$2')
 
plt.grid(True)
plt.legend()
plt.ylabel('p(x)')
plt.xlabel('x')
plt.title('Gamma Distribution')
plt.savefig('3.Gamma Distribution.jpeg')
cs

 

 

 

 

 

Original formula or English plot

Scientific clarifications, separate from the historical source

The 2017 author’s statements above are translated literally, including historical errors.

λ is the positive arrival rate per unit time, and E[N(t)]=λt. Waiting for the α-th event and summing α exponential variables requires an integer α≥1. The Gamma density extends to every real α>0; this does not mean a fractional event count.

The finite sum for fewer than α events is survival P(T>t); its complement is the CDF P(T≤t). The historical exact-time probability statements in P111/P113/P114 and img_121 should be read as density descriptions: for a continuous T, P(T=t)=0. Interval probabilities are obtained by integrating f. The historical symbol P(t) in img_121 and img_122 labels a density here; it is not the probability of an exact waiting time.

Γ(a) is defined by its integral for real a>0; Γ(n)=(n−1)! is the positive-integer specialization. In the source convention α is shape, λ is rate and θ=1/λ is scale. The Gamma function in P130 is distinct from the Gamma probability distribution. P117 literally says “alpha and gamma”; surrounding formulas use alpha and lambda. In P128, μ is the Gaussian mean, σ the standard deviation, and σ² the variance.

The MGF is finite for its argument t<λ and diverges at t≥λ. The historical sheets rename the density variable to x while using t for the MGF argument. The mean and variance are α/λ and α/λ². Independent Gamma variables add their shapes when they share rate λ.

The source code uses a=alpha and scale=1/llambda with the default location zero. Its loc=alpha/llambda assignment is not passed to gamma.pdf. The five plotted (α,λ) pairs are (1,.5),(2,.5),(3,.5),(5,1),(9,2). Unused imports and assignments are preserved.

The clipped historical img_054 right/bottom strokes and img_120 margin-note endings remain unknown. The legible continuation and cleaner derivative sheet are independent source evidence; they do not recover missing handwriting.

Original source

Derivation of the Gamma Distribution — original Naver post by gdpark, November 16, 2017

Comments

Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.