Normal hypothesis tests for a population mean
A Normal hypothesis test for a mean uses an observation or sample mean to assess a claim about an unknown population mean . The calculation asks how unusual the observed result would be if the null hypothesis were true.
The key model for a random sample of size is
Therefore the standard deviation of , called its standard error, is
This lesson concerns A level questions in which the population variance is known. It does not cover the test used when must be estimated.
Prerequisites
Section titled “Prerequisites”You should be able to:
- standardise a Normal random variable using ;
- use inverse Normal probabilities and calculator distribution functions;
- distinguish a population mean from a sample mean ;
- write null and alternative hypotheses and interpret a significance level.
Review the Normal distribution and hypothesis testing language where needed.
From a population to a test statistic
Section titled “From a population to a test statistic”Suppose the null hypothesis is . Under ,
so an observed sample mean can be standardised as
The sign and magnitude of both matter:
- means that is above the null mean;
- means that is below the null mean;
- measures the distance from the null mean in standard errors.
For one observation, , so the denominator is simply .
Worked example 1: form the null model
Section titled “Worked example 1: form the null model”Packet masses are modelled by grams. A random sample of packets is taken. Under the claim that the population mean remains g,
Thus g. Sample means vary much less than individual packet masses.
If the sample mean is g, then
The sample mean is standard errors below the claimed mean.
Choose the alternative hypothesis first
Section titled “Choose the alternative hypothesis first”Let be the population mean and let be the value claimed under .
| Investigation | Hypotheses | Relevant evidence |
|---|---|---|
| Has the mean increased? | , | upper tail |
| Has the mean decreased? | , | lower tail |
| Has the mean changed? | , | both tails |
Hypotheses concern the population parameter , not the observed statistic . Decide whether the test has one tail or two from the question, before inspecting the data.
Self-check 1
Section titled “Self-check 1”A filling machine is set to deliver a mean of ml. An engineer investigates whether it is overfilling. A sample has mean ml. Write the hypotheses.
Answer
Let be the population mean fill volume. Since overfilling means a larger mean,
Writing a hypothesis about is incorrect because ml is already observed.
Method 1: use a p value
Section titled “Method 1: use a p value”The p value is the probability, assuming , of obtaining the observed result or one more extreme in the direction specified by .
For an observed standardised value :
Then
For a two tailed test, doubling the smaller tail probability counts equally extreme results on the opposite side of . This works because the Normal distribution is symmetric.
Worked example 2: lower tailed test
Section titled “Worked example 2: lower tailed test”A manufacturer states that the lifetime of a component is Normally distributed with standard deviation hours and mean hours. A random sample of components has mean lifetime hours. Test at the level whether the population mean lifetime has decreased.
Let be the population mean lifetime. The hypotheses are
Under ,
so the standard error is
The observed test statistic is
This is a lower tailed test, so
Since , reject . There is sufficient evidence at the significance level to suggest that the population mean lifetime has decreased.
Worked example 3: upper tailed test with no rejection
Section titled “Worked example 3: upper tailed test with no rejection”The mass of produce in a container is modelled as Normal with known standard deviation g. Historically its mean is g. After a process adjustment, a random sample of containers has mean mass g. Test at the level whether the mean has increased.
Under , the standard error is
Therefore
and the upper tail probability is
Since , do not reject . There is insufficient evidence at the level to suggest that the population mean mass has increased.
The sample mean is above g, but the difference is not large relative to ordinary sampling variation.
Two tailed tests
Section titled “Two tailed tests”A two tailed alternative treats a result equally far below or above as evidence against . At overall significance level , each tail contains probability .
Worked example 4: calculate and double one tail
Section titled “Worked example 4: calculate and double one tail”A machine produces rods whose lengths are Normally distributed with known standard deviation mm. It is set to a mean of mm. A random sample of rods has mean mm. Test at the level whether the mean length has changed.
The standard error is
giving
The upper tail probability is approximately
Hence the two tailed p value is
Since , reject . There is sufficient evidence at the level to suggest that the population mean rod length has changed.
Self-check 2
Section titled “Self-check 2”A two tailed test gives . Find the p value and state the decision at the level.
Answer
By symmetry,
Since , do not reject .
Method 2: use critical values
Section titled “Method 2: use critical values”A critical value separates results that lead to rejection from those that do not. Common standard Normal boundaries are:
| Test | Significance level | Reject when |
|---|---|---|
| upper tailed | ||
| lower tailed | ||
| two tailed | $ | |
| upper tailed | ||
| lower tailed | ||
| two tailed | $ |
You may instead convert a boundary into a critical sample mean. For an upper tailed test,
For a lower tailed test, use the negative boundary. For a two tailed test, use in both directions.
Because the Normal model is continuous, whether a single boundary point is written with or does not alter its probability. The usual decision convention includes equality in the critical region.
Worked example 5: critical sample mean
Section titled “Worked example 5: critical sample mean”Weekly demand is Normally distributed with known standard deviation units. A supplier tests
using the mean of weeks at the level. Find the critical value of .
Under , the standard error is
The upper standard Normal boundary is , so
Thus the critical region is approximately
A sample mean of lies in this region, so would be rejected. A sample mean of does not.
Worked example 6: two critical boundaries
Section titled “Worked example 6: two critical boundaries”The fill mass of a product is Normal with known standard deviation g. A sample of items is used to test against at the level.
The standard error is
For a two tailed test, the boundary is . Therefore
The critical regions are
Notice that is placed in each tail, giving in total.
Testing a single observation
Section titled “Testing a single observation”Some questions give one observation from a Normal population rather than a sample mean. Then use :
Worked example 7: one observation
Section titled “Worked example 7: one observation”Under a stated model, journey time is Normal with standard deviation minutes and mean minutes. One randomly selected journey takes minutes. Test at the level whether journey times are longer than the model claims.
For one observation,
Thus
Reject . There is sufficient evidence at the level to suggest that the population mean journey time is greater than minutes.
Do not divide by another square root: the observation is not a sample mean based on an unstated sample size.
A reliable exam method
Section titled “A reliable exam method”- Define in context.
- Write and , choosing the tail from the question.
- Under , write the distribution of or .
- Calculate the correct standard deviation, using for a sample mean.
- Find a p value or critical region.
- Compare with the significance level and state reject or do not reject .
- Give a contextual conclusion using the direction in .
Worked example 8: full test from start to finish
Section titled “Worked example 8: full test from start to finish”The breaking strength of a cable is Normally distributed with standard deviation N. The manufacturer claims a mean strength of N. A regulator suspects the true mean is lower. A random sample of cables has mean strength N. Test the regulator’s suspicion at the level.
Let be the population mean breaking strength in newtons.
Under ,
with standard error
Hence
The lower tail p value is
Since , do not reject . There is insufficient evidence at the significance level to suggest that the population mean breaking strength is below N.
The result would be significant at , but the question demands the stricter standard.
Common misconceptions
Section titled “Common misconceptions”- Using instead of . A sample mean has less variation than an individual value.
- Using variance in the denominator. Standardisation always divides by a standard deviation.
- Writing hypotheses about . The hypotheses concern the unknown population mean .
- Choosing the tail from the data. The wording of the investigation determines .
- Forgetting the second tail. For , double the smaller one tail probability.
- Saying that is proved. A non-significant result means only that the evidence was insufficient to reject .
- Ignoring context. Finish with a statement about the population mean, not only a calculator probability.
Final self-check
Section titled “Final self-check”A process produces values that are Normally distributed with known variance . A random sample of values has mean . Test against at the level.
Answer
Since , the standard error is
Therefore
and
Since , reject . There is sufficient evidence at the level to suggest that the population mean is less than .
Using the critical value method gives the same decision because .
Next steps
Section titled “Next steps”- Compare exact discrete testing in binomial hypothesis tests.
- Learn how a sample association is tested in hypothesis tests for correlation.
- Revisit sampling to understand why random and independent observations matter.
- Practise selecting an appropriate model in choosing a distribution.