Binomial hypothesis tests: critical regions and significance
A binomial hypothesis test uses sample data to assess a claim about the probability of success in a repeated trial. It asks whether the observed number of successes would be sufficiently unusual if a stated value of were true.
The test does not prove that a hypothesis is true or false. It makes a controlled decision from uncertain evidence.
Prerequisites
Section titled “Prerequisites”You should be able to:
- recognise when is an appropriate model;
- calculate individual and cumulative binomial probabilities;
- interpret , and complements;
- understand the language of hypothesis testing.
Review the binomial distribution if you are not yet confident with cumulative probabilities.
The central idea
Section titled “The central idea”Suppose a coin is claimed to have probability of landing heads. It is tossed times and gives heads. If the claim is true, then
where is the number of heads.
Seventeen heads is possible under this model, but possibility is not the issue. We ask:
How probable is a result at least as extreme as the one observed, assuming the claim is true?
If that probability is at most the chosen significance level, usually , the result is called significant and the null hypothesis is rejected.
Hypotheses and tails
Section titled “Hypotheses and tails”The null hypothesis gives the value of used in the probability calculation. The alternative hypothesis expresses the departure being investigated.
| Investigation | Null hypothesis | Alternative hypothesis | Test |
|---|---|---|---|
| Has increased? | upper tailed | ||
| Has decreased? | lower tailed | ||
| Has changed? | two tailed |
The wording of the claim determines the alternative hypothesis. Write the hypotheses before looking at whether the observed result is high or low.
Worked example 1: choosing the hypotheses
Section titled “Worked example 1: choosing the hypotheses”A seed supplier claims that of a variety germinate. A gardener suspects that the germination rate is lower.
Let be the probability that a randomly selected seed germinates. The suspicion is specifically about a decrease, so
This is a lower tailed test.
If the question had asked whether the advertised rate was incorrect, the alternative would instead be .
The significance level
Section titled “The significance level”The significance level is the largest intended probability of rejecting when is actually true. Common choices are
A test demands stronger evidence than a test. A result can therefore be significant at but not at .
Rejecting a true null hypothesis is a Type I error. The significance level controls the probability of that error. Because a binomial distribution is discrete, the actual probability of rejection is often less than, rather than exactly equal to, .
Method 1: compare a tail probability with
Section titled “Method 1: compare a tail probability with α\alphaα”For an observed value :
- upper tailed test: calculate ;
- lower tailed test: calculate ;
- two tailed test: use both tails, with the significance level shared between them.
The tail probability for a one tailed test is the probability, under , of the observed result or one more extreme in the direction of .
Then apply the decision rule
Equality leads to rejection because the critical region is constructed to have probability at most .
Worked example 2: an upper tailed test
Section titled “Worked example 2: an upper tailed test”A machine normally produces acceptable components with probability . After servicing, randomly selected components are inspected and are acceptable. Test at the significance level whether the probability of an acceptable component has increased.
Let be the probability that a component is acceptable. Then
Under ,
The alternative points towards unusually large values, so calculate
Since
reject . There is sufficient evidence at the significance level to suggest that the probability of an acceptable component has increased.
Worked example 3: a lower tailed test
Section titled “Worked example 3: a lower tailed test”A player has historically made a free throw with probability . In a training session, she makes of attempts. Test at the level whether her success probability has decreased.
Let be the probability that she makes a free throw.
Under ,
Since is in the direction of ,
As , do not reject . There is insufficient evidence at the significance level to suggest that her free throw probability has decreased.
This does not establish that . The sample has simply failed to provide sufficiently strong evidence against that value.
Self-check 1
Section titled “Self-check 1”A website has historically converted of visits into sales. After a redesign, sales occur in independent visits. A test of against gives
State the conclusion at the significance level.
Answer
Since , do not reject . There is insufficient evidence at the level to suggest that the redesign has increased the conversion probability.
The observed rate is greater than , but the difference is not statistically significant at the stated level.
Critical regions
Section titled “Critical regions”A critical region is the set of values of that cause to be rejected. Its boundary value is a critical value.
For an upper tailed test, the critical region has the form
where is the smallest integer satisfying
For a lower tailed test, it has the form
where is the largest integer satisfying
You must check the neighbouring value to show that the region is as large as possible without exceeding the significance level.
Worked example 4: upper critical region
Section titled “Worked example 4: upper critical region”Under , . Find the upper critical region for a test.
Calculator values give
and
The region is too probable because . The next possible region is within the limit because . Therefore
is the critical region, with critical value .
Its actual significance level is
It is not . No integer boundary produces every probability between and .
Worked example 5: lower critical region
Section titled “Worked example 5: lower critical region”Under , . Find the lower critical region for a test.
We seek the largest for which . Suppose the cumulative probabilities are
Including would make the probability exceed , so
is the critical region. Its actual significance level is .
If the observed value were , it would lie just outside the critical region, so would not be rejected at .
Self-check 2
Section titled “Self-check 2”Under , . For a lower tailed test,
State the critical region and its actual significance level.
Answer
The critical region is
Its actual significance level is
The value cannot be included because .
Two tailed tests
Section titled “Two tailed tests”A two tailed test uses
Both unusually small and unusually large observations count as evidence against . At overall significance level , A level questions normally place at most in each tail.
For a two tailed test, find:
- a lower region with probability at most ;
- an upper region with probability at most .
The two parts need not have equal actual probabilities. A binomial distribution is discrete and may be asymmetric.
Worked example 6: construct a two tailed critical region
Section titled “Worked example 6: construct a two tailed critical region”A coin is claimed to land heads with probability . It is tossed times. Find the critical region for a two tailed test.
Under ,
Each tail may contain at most . For the lower tail,
Thus the lower part is .
The distribution is symmetric because , so
while . Hence the upper part is .
The critical region is
The actual significance level is the probability of the whole critical region:
If heads were observed, reject . There would be sufficient evidence that the probability of heads differs from .
Worked example 7: an asymmetric two tailed test
Section titled “Worked example 7: an asymmetric two tailed test”For under , find the critical region for a two tailed test from the following values:
Each tail is allowed at most .
For the lower tail, is allowed because . For the upper tail, is not allowed because , so the boundary must move to . Therefore
The actual significance level is approximately
The tail probabilities are unequal. Trying to force equal probabilities would discard valid outcomes and make the test unnecessarily conservative.
Calculator strategy
Section titled “Calculator strategy”Calculator menus differ, but the probability identities do not.
For :
and
The shift to matters. For example,
not .
For an isolated value,
Keep unrounded calculator values when comparing with . Round only when presenting the probability. A displayed value of might lie slightly above or below before rounding.
A complete exam method
Section titled “A complete exam method”- Define in context.
- State and .
- Under , write .
- Identify whether the test is upper tailed, lower tailed or two tailed.
- Calculate the relevant tail probability, or construct the critical region.
- Compare with the significance level, or check whether the observation lies in the critical region.
- State “reject ” or “do not reject ”.
- Give a conclusion in the context of the question.
Worked example 8: full test using a critical region
Section titled “Worked example 8: full test using a critical region”A manufacturer claims that of its batteries are defective. A customer believes the defect probability is higher. In a random sample of batteries, are defective. At the significance level, test the customer’s belief.
Let be the probability that a randomly selected battery is defective.
Under ,
This is upper tailed. The relevant probabilities are
and
Therefore the critical region begins at :
The observed value lies in the critical region, so reject . There is sufficient evidence at the significance level to support the customer’s belief that the probability of a defective battery is greater than .
The result is only just significant. That does not change the decision, but it is useful context when interpreting the strength of evidence.
Assumptions behind the test
Section titled “Assumptions behind the test”The test is valid only if a binomial model is reasonable. Check that:
- there is a fixed number of trials;
- each trial has two outcomes, labelled success and failure;
- trials are independent;
- the probability of success is constant across trials.
In sampling without replacement, independence is only approximate. It is generally reasonable when the population is much larger than the sample. Bias in the selection process is not repaired by a hypothesis test.
Worked example 9: challenge the model
Section titled “Worked example 9: challenge the model”A school surveys the first pupils entering the library and treats “supports the proposal” as success. It then uses a binomial test to make a claim about all pupils.
Even if each response has two outcomes, the sample may be biased because library users might differ systematically from the whole school. The calculation could be arithmetically correct while the conclusion is unreliable.
Statistical inference depends on how the data were obtained, not only on the formula used afterwards. Review sampling methods and probability modelling.
Common misconceptions
Section titled “Common misconceptions””Not significant” means the null hypothesis is true
Section titled “”Not significant” means the null hypothesis is true”No. It means the evidence was not strong enough to reject at the chosen level. A larger sample might detect a real difference that this sample did not.
The observed value alone is the probability
Section titled “The observed value alone is the probability”For an upper tailed observation , use , not . Outcomes more extreme than also count as evidence against .
The sample proportion replaces under
Section titled “The sample proportion replaces ppp under H0H_0H0”If successes occur in trials, the sample proportion is . But if , probabilities for the test are calculated using , not .
The tail follows the data
Section titled “The tail follows the data”The research question chooses the tail. If the stated alternative is , an unexpectedly low result does not justify changing to a lower tailed test after seeing the data.
A critical region must have probability exactly
Section titled “A 5%5\%5% critical region must have probability exactly 0.050.050.05”Binomial values are discrete. The actual significance level is the largest available critical probability not exceeding the stated level, and is often smaller.
Mixed self-check
Section titled “Mixed self-check”Question 1
Section titled “Question 1”A die is suspected of producing sixes more often than a fair die. In rolls, it produces sixes.
- State suitable hypotheses.
- Write the null distribution.
- Given that , complete a test.
Answer
Let be the probability of rolling a six.
Under ,
Since , reject . There is sufficient evidence at the level to suggest that the die produces sixes with probability greater than .
Question 2
Section titled “Question 2”For a lower tailed test, under . The following probabilities are given:
Find the critical region. State the decision if successes are observed.
Answer
The critical region is
with actual significance level . Since is outside the critical region, do not reject .
Question 3
Section titled “Question 3”Explain why a two tailed test does not usually place in each tail.
Answer
Placing in each tail would give a total intended significance level of up to . The overall allowance is shared, normally as at most in each tail.
Question 4
Section titled “Question 4”A test produces a p value of . State the decisions at the and significance levels.
Answer
Since
reject at both levels. The result is significant at both and .
Summary
Section titled “Summary”For under :
Reject when the relevant probability is at most , or when the observation lies in the critical region. Always finish with a cautious conclusion in context.
Next, compare this exact discrete procedure with normal hypothesis tests and apply the same decision language in correlation hypothesis tests.