Hypothesis Testing
An introductory walkthrough of hypothesis tests, significance levels, p-values, power, and normal, t, and chi-square examples.
English translation of the original Statistics #8 post. The author’s original wording, notation, and claims are preserved. Separately labeled editorial notes explain source errors and statistical qualifications.
I will just cover some very basic things about testing and move on.
First, what is hypothesis testing?
“When we judge an unknown parameter from observed samples,
we want to manage the possibility of error in that judgment at a level fixed beforehand.” Apparently, that is what hypothesis testing means.
Then let me just throw out an example first.
EX. Suppose that the lifespans of lightbulbs produced by a company can be assumed to follow a normal distribution,
and quality control keeps the mean at 1500 hours and the standard deviation at 100 hours.
That is, if we represent a bulb’s lifespan by the random variable X, X follows M(1000, 100²).
Now this company’s research team claims to have developed a new product with the same cost and standard deviation as the existing product, but a longer mean lifespan.
To check this, they test-produced 25 bulbs and measured their mean lifespan,
obtaining
.
From this result, can we be sure that the new bulb’s mean lifespan has become longer than the existing bulb’s mean lifespan?
When we are faced with a question like this,
we carry out hypothesis testing.
First, since the research team is claiming that the mean lifespan is 1500 hours or more, let us set up the hypotheses like this.
θ represents the “mean lifespan.” I will explain hypothesis testing little by little by explaining what H₀ and H₁ are.
These folks’ claim was that “the bulb’s mean lifespan is longer than the existing mean lifespan of 1500,”
and that claim is written there as
.
This
is called the alternative hypothesis.
And, written as
, this
is called the null hypothesis.
Here, “null” is the null in “nullify”: to make something invalid.
That is, “
is the hypothesis that nullifies the alternative hypothesis
.”
We set up these two hypotheses, and the act of deciding which hypothesis to adopt and which to reject is called “hypothesis testing.”
But we need some basis and a convincing enough argument to shut everyone up before we can claim that we adopt one hypothesis and reject the other.
There are various criteria for testing, but here we will use an example with the Z-statistic.
That is because there are still more basic terms to sort out.
Let us think about the Z-statistic.
The standard deviation of
is
, so it is 20.
And under H₀,
.
When
is large enough (when the mean of the things we sampled for the trial comes out absurdly large), that is, when
,
we will reject
and adopt the alternative hypothesis
.
Why is ![]()
the particular cutoff? Well, that would be up to the person doing the test....
But in statistics, apparently
or
are used a lot. (I will mention this again later.)
Anyway,
if our measured sample mean exceeds that value, we can decide to reject H₀ and adopt H₁—that is the idea!
The deeper meaning of this will come up again later!!!
Before that, let us organize the terminology here.
A statistic that serves as the criterion for testing, such as the Z-statistic when we test hypotheses using a criterion like this,
is called a test statistic.
An area like Z ≥ 1.645, the “area where the null hypothesis is rejected,” is called the critical region.
The remaining area is called the acceptance region.
(cf. The boundary is called the critical point.)
(The criterion is whether we “reject” the null hypothesis or “adopt” the alternative hypothesis.)
If we draw the critical region and the acceptance region,
Untouched original source image

then let us look a little more closely at what Z ≥ 1.645 means.
First, there are two kinds of “errors” in testing:
1. When the null hypothesis H₀ is true, adopting the alternative hypothesis H₁: a Type I error.
2. When the alternative hypothesis H₁ is true, adopting the null hypothesis H₀: a Type II error.
So,
| H₀ true | H₁ true | |
|---|---|---|
| Adopt H₀ | Good decision | Type II error |
| Adopt H₁ | Type I error | Good decision |
Untouched original source image

suppose we observe the sample mean X̄ and look at the Z-statistic calculated from it.
It falls in the critical region, so I reject the null hypothesis.
But,
Untouched original source image

the original meaning of this graph is “the probability of being drawn when we draw X̄,” right?
That is, “the probability of drawing an X̄ that falls in the critical region”—the green part over there in the picture!!!
So what Z ≥ 1.645 means is “the probability that, even though the null hypothesis is true, X̄ happens to come out large in that trial and we reject H₀.”
In other words, “the probability of committing a Type I error” is this
Untouched original source image

.
We say that this number means “the permissible limit of the probability of the error of rejecting H₀ when it is true and adopting H₁ instead.”
It is called the significance level, and apparently it is commonly written as α.
In other words, α: 5% means that a true hypothesis is not rejected four or more times out of 100.
We can say that the confidence level is 95%, and the confidence coefficient is 1 − α—that is what it means!
For reference, the significance levels commonly used are 1%, 5%, and 10%, apparently.
Anyway, the research team’s sample mean from a sample of size 25 was 1550,
so Z = 2.5,
and Z = 2.5 is greater than 1.625 (it falls in the critical region).
So “at significance level 5%,” we can conclude that we reject the null hypothesis and adopt the alternative.
At this point, let us look at the p-value.
The mean the research team obtained from its sample was 1550, and the corresponding Z-statistic
was Z = 2.5, right? Then, if Z = 2.5,
is what that says.
“We would have rejected H₀ even if we had tested at significance level 0.6%,
and the significance level would have had to be below 0.6% for us not to reject H₀”.....
This 0.6% is called the p-value, apparently (or the significance probability).
The meaning of the p-value is “the smallest significance level at which we can reject H₀ for the test statistic.”
So, when X̄ is observed as 1550 and when it is observed as 1600,
we would reject H₀ in both cases,
but the larger X̄ is, the smaller the p-value becomes. The p-value represents how the degree of certainty that H₁ is true can differ between the two cases.
Now those terms that felt so unfamiliar are gradually starting to feel familiar.
Let me lightly organize the terms we have not mentioned yet:
A hypothesis like
, where the parameter value is specified as a single point, is a simple hypothesis.
A hypothesis like
, where the parameter value is specified as a range, is called a composite hypothesis... ;;; just terminology ;;;
But apparently the null hypothesis is generally expressed as a simple hypothesis and the alternative as a composite hypothesis.
Power function
1. In a testing method, the function expressing the probability of rejecting H₀ (the probability of entering the critical region) as a function of the parameter θ is called the power function.
Mathematically,
it is expressed as
Untouched original source image

!!
2. The value of the power function at a particular θ belonging to H₁ is called the power, apparently.
Suppose
. And when we use a sample of n = 25,
let us suppose
.
follows the standard normal distribution, and if we calculate π(θ) for a given θ,
Untouched original source image

(Here uppercase Φ would be the “cumulative probability,” riiight?)
Then, for example, if we keep calculating π(θ) for θ = 0, θ = 0.4, θ = 1, and so on and on,
we get some values.
These values come out...
and the graph of them is what we call the power function.
Here, let us think only about that marked point.
We have set the value of π(θ) at θ = 0 to 0.025.
That means the probability of rejecting H₀ when θ = 0.
And the probability of rejecting it
Untouched original source image

means the probability here,
so we have seen this somewhere before, have we not?!?!?!?
It was the picture talking about the significance level, right???
So the meaning of the value of π(θ) at each θ in the power function
was the α value!!! Another name for “the probability of committing a Type I error!”
Now shall we look at this with the t-statistic?
(And while we are at it, we looked only at one-sided examples before, so let us look at the meaning of two-sided testing here too.)
As we saw earlier, when X
follows
,
it follows
.
So
follows the standard normal distribution, which is nice,
but usually we do not know
. That is the catch.
Suppose we replace the
there with
.
That is, we have replaced it with
.
Now it no longer follows a normal distribution. As we saw earlier, we have to regard it as following a t-distribution with n − 1 degrees of freedom.
Then this time, let us look at the 95% confidence interval and the critical t-value using the T-statistic!
With a “two-sided test”!!! (We will use 27 degrees of freedom. That means the sample size is 28.)
If we consult the t-distribution table,
we can confirm
.
Because the table
Untouched original source image

tells us that!
That is, when the degrees of freedom are 27, the probability that the t-value falls in (−2.052, 2.052) is 95%.
Here, −2.052 and 2.052 are called the critical t-values.
t = −2.052 is called the lower critical t-value, and the other is called the upper critical t-value.
Now, if we change ![]()
to this, it means that the probability that μₓ lies in this interval is 95%.
That is, when μₓ is observed, we can expect it to be in that interval 95 times out of 100,
and the significance level is 5%, or the probability of a Type I error is 5%... exactly what we said before.
This time, I will introduce a simple example of testing with the chi-square distribution and head out.
Testing with the chi-square distribution is generally for when we know S but test an unknown σ.
The basic idea is the same as the testing we dealt with earlier, so it should not be difficult.
Suppose that S² for 31 observations from a normally distributed population is 12.
And
let us set up
.
Then the statistic
follows the distribution
.
And if we take α = 5%, the confidence interval looks like this.
But if σ² is 9, then
, so it lies in the acceptance region in the graph above.
Therefore, “we do not reject H₀” will be the conclusion.
That was a surface-level look at testing.
After this, I will deal with more detailed problems in “statistics” posts rather than “statistics basics” posts.
Comments
Discussion happens via GitHub Discussions. You'll need a GitHub account to comment.