Moment Generating Function

Moment generating functions: raw moments, uniqueness of a distribution, and the product rule for sums of independent random variables.

Moment Generating Function (moment generating function)


When we get some bundle of data, we can calculate things like what the mean is and what the variance is.

We are doing it to get a good bit of the data’s characteristics out of it, so, well.....

The idea of calculating the mean, the variance, and so on does not feel terribly objectionable.


But actually, another name for the ‘mean’ is the first moment (1st moment).

And the ‘variance’ is the second moment (2nd moment).


As for which numbers those 1 and 2 are connected to,

the definition of the mean is \(\mu = \sum x_i^{\color{red}{1}}\cdot f(x_i)\) like this, and it is connected to the ‘1’ that I have colored red up there,

and the definition of variance is \(\sigma^2 = \sum (x_i-\mu)^{\color{red}{2}}\cdot f(x_i)\) like this, and it turns out it was connected to the ‘2’ that I have colored red up there.



Now, so far I have only talked up to the second moment, and you learn this much in high school too,

but if there is a 1 and there is a 2, what about 3? 4? n????????????



Yes! They exist

\[\mathit{skewness}(\text{skewness}) = \sum (x_i-\mu)^{\color{red}{3}}\cdot f(x_i)\]


\[\mathit{kurtosis}(\text{kurtosis}) = \sum (x_i-\mu)^{\color{red}{4}}\cdot f(x_i)\]


.

.

.


\[n\text{-}th\ moment = \sum (x_i-\mu)^{\color{red}{n}}\cdot f(x_i)\]






This is the ‘moment’ part of moment generating function, right??

As you can see from the name, mathematicians came up with the moment generating function so they could get moments out well, whenever and in whatever situation.

The moment generating function is written as \(\mathrm{M}_X(t)\) and its definition is \(\mathrm{M}_X(t) = \mathrm{E}(e^{tX})\) .


Why is it defined this way??? It is very simple.

\(\mathrm{E}(e^{tX})\)   is, by the definition of expected value,


\[\begin{aligned}\mathrm{E}(e^{tX}) &= \sum e^{tX_i}f(X_i)\\ &= e^{tX_1}f(X_1)+e^{tX_2}f(X_2)+e^{tX_3}f(X_3)+\cdots\end{aligned}\]


That is,

\(\mathrm{M}_X(t) = e^{tX_1}f(X_1)+e^{tX_2}f(X_2)+e^{tX_3}f(X_3)+\cdots\) that is what it is.



If we differentiate both sides with respect to t,

\[\mathrm{M}_X^{\prime}(t) = X_1 e^{tX_1}f(X_1)+X_2 e^{tX_2}f(X_2)+X_3 e^{tX_3}f(X_3)+\cdots\]



if we differentiate both sides of this with respect to t one more time,

\[\mathrm{M}_X^{\prime\prime}(t) = X_1^2 e^{tX_1}f(X_1)+X_2^2 e^{tX_2}f(X_2)+X_3^2 e^{tX_3}f(X_3)+\cdots\]



shall we put t = 0 into each of the two expressions??????????


\[\mathrm{M}_X^{\prime}(0) = X_1 e^{0\cdot X_1}f(X_1)+X_2 e^{0\cdot X_2}f(X_2)+X_3 e^{0\cdot X_3}f(X_3)+\cdots = \sum X_i\cdot f(X_i)\]


\[\mathrm{M}_X^{\prime\prime}(t) = X_1^2 e^{0\cdot X_1}f(X_1)+X_2^2 e^{0\cdot X_2}f(X_2)+X_3^2 e^{0\cdot X_3}f(X_3)+\cdots = \sum X_i^2 f(X_i)\]



An expression with these characteristics is the moment generating function

so, to sum things up here before moving on,

\(\mathrm{M}_X^{(m)}(0) = \mathrm{E}(X^m)\)  it is because of this property!!!!





But the real reason I am going over the moment generating function here

is because of a property of the moment generating function that we will use later.


As for what property we will use,


suppose there are two random variables called X and Y, and the two random variables have f(X) and g(Y), respectively, as their probability distributions,

and if X and Y have the same space,


→ they say that the moment generating functions being equal means, precisely, that f=g.


In other words, if the mgf (moment generating function) exists,

we can say that there is just one probability distribution corresponding to the mgf!!!!!!!!!!!!






Also, as a little side branch,

let us say there are mutually independent random variables X, Y, and Z.

Then their respective moment generating functions would be written \(\mathrm{M}_X(t)\), \(\mathrm{M}_Y(t)\), \(\mathrm{M}_Z(t)\) like this, right??????


But in this case

the moment generating function of the random variable X+Y+Z is \(\mathrm{M}_{X+Y+Z}(t) = \mathrm{M}_X(t)\mathrm{M}_Y(t)\mathrm{M}_Z(t)\) that is the fact.


I used three of them as an example, but it is possible for n of them too.


\[\mathrm{M}_{X+Y+Z}(t) = \mathrm{E}(e^{t(X+Y+Z)}) = \mathrm{E}(e^{tX}\cdot e^{tY}\cdot e^{tZ}) = \mathrm{E}(e^{tX})\cdot\mathrm{E}(e^{tY})\cdot\mathrm{E}(e^{tZ}) = \mathrm{M}_X(t)\mathrm{M}_Y(t)\mathrm{M}_Z(t)\] 

Well, it is because of this sort of reasoning.



Since this moment generating function gets used a little later, I have briefly gone over it first before moving on!

Then next time I will start with the story of ‘distributions’!

Scientific clarifications to the historical learning note.

The historical prose and formulas above remain unchanged.

Raw moments and central moments

The mean is the first raw moment, $\mu=\mathrm{E}(X)$. The $r$-th raw moment is $\mathrm{E}(X^r)$, whereas the $r$-th central moment is $\mathrm{E}[(X-\mu)^r]$. The first central moment is zero, and variance is the second central moment. Derivatives of the MGF at zero give raw moments, so $\operatorname{Var}(X)=M_X^{\prime\prime}(0)-[M_X^{\prime}(0)]^2$. The displayed general formula involving $(x_i-\mu)^n$ is a central-moment formula.

Standardized skewness and kurtosis

The historical skewness and kurtosis formulas show the third and fourth central moments without standardization. When $\sigma>0$ and the required moments are finite, the usual dimensionless coefficients are $\gamma_1=\mathrm{E}[(X-\mu)^3]/\sigma^3$ and $\beta_2=\mathrm{E}[(X-\mu)^4]/\sigma^4$. Excess kurtosis is $\beta_2-3$. These are distinct from the unstandardized quantities printed above.

The derivative evaluated at zero

In the second-derivative line with $e^{0\cdot X_i}$ on the right, the historical left side still says $M_X^{\prime\prime}(t)$. After substituting $t=0$, the consistent left side is $M_X^{\prime\prime}(0)$, and the result is $\mathrm{E}(X^2)$, the second raw moment.

When the MGF generates moments

The usual differentiation result assumes that $M_X(t)=\mathrm{E}(e^{tX})$ is finite on an open interval containing zero. Under this condition, $M_X^{(r)}(0)=\mathrm{E}(X^r)$ for every nonnegative integer $r$. For a discrete law, $f(x_i)$ in the sums denotes probability masses whose sum is one; a continuous law uses an integral against its distribution instead. A value at zero alone is insufficient: every probability law has $M_X(0)=1$.

What the uniqueness statement requires

If two MGFs are finite and equal throughout an open interval containing zero, their random variables have the same probability distribution. They do not need to be defined on the same probability space. In the discrete notation above, equality of the laws means equality of their probability masses; continuous density representatives are determined only up to sets of measure zero. Equality only at $t=0$ does not establish uniqueness.

The product rule and independence

For mutually independent random variables, $M_{X+Y+Z}(t)=M_X(t)M_Y(t)M_Z(t)$ at values of $t$ where the MGFs are finite. The factorization follows from independence of $e^{tX}$, $e^{tY}$ and $e^{tZ}$. The same argument applies to any finite number of mutually independent variables; pairwise independence alone does not generally suffice.

Original Korean learning note

Comments

Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.