The normal distribution
The normal distribution is a continuous probability model with a symmetric, bell-shaped density curve. It is used for variables whose values cluster around a central value, with increasingly extreme values becoming increasingly rare.
If a continuous random variable has mean and variance , write
The second parameter is the variance, not the standard deviation. Thus
means that and .
This lesson explains the model, probability calculations, inverse problems, parameter estimation and the normal approximation to a binomial distribution.
Prerequisites
Section titled “Prerequisites”You should be able to:
- use complements and combine intervals from probability;
- interpret the mean, variance and standard deviation from averages and spread;
- distinguish continuous and discrete random variables;
- use the binomial distribution;
- solve simultaneous equations and use inverse calculator functions.
Shape and parameters
Section titled “Shape and parameters”Every normal density curve has these properties:
- it is symmetric about ;
- its mean, median and mode are all ;
- the total area under the curve is ;
- the curve approaches, but never meets, the horizontal axis;
- its points of inflection are at and .
The mean controls the location of the curve. Increasing translates the curve to the right without changing its shape.
The standard deviation controls the spread. A larger produces a wider, lower curve because the total area must remain .
The density function is
but A-level probability calculations use calculator distribution functions rather than direct integration of this expression.
The empirical rule
Section titled “The empirical rule”For every normal distribution, approximately
These values are useful for judging whether a calculator answer is plausible. They should not replace a calculator when an accurate probability is required.
Worked example 1: reading the parameters
Section titled “Worked example 1: reading the parameters”The masses grams of packets are modelled by
State the mean and standard deviation, and estimate the central interval containing about of packet masses.
The mean is
Since ,
About of values lie within two standard deviations of the mean:
Therefore the estimated interval is
Self-check 1
Section titled “Self-check 1”Suppose . State the axis of symmetry and the coordinates of the two points of inflection.
Answer
Here and . The axis of symmetry is
and the points of inflection occur when , so their coordinates are
Continuous probabilities
Section titled “Continuous probabilities”Since a single point has zero width,
for every continuous random variable. Consequently,
and similarly
This differs fundamentally from a discrete distribution, where may be positive.
Normal distribution calculators usually offer:
- normal CDF or normal cumulative probability for an area between bounds;
- inverse normal for a boundary corresponding to a given cumulative area.
Always identify the required region before entering values. For an upper tail, either use an upper bound of or calculate a complement.
Worked example 2: a lower-tail probability
Section titled “Worked example 2: a lower-tail probability”The lifetime hours of a component is modelled by
Find the probability that a component lasts fewer than hours.
The required area is to the left of :
Using normal CDF with lower bound , upper bound , mean and standard deviation gives
so
The answer is below , as expected because is below the mean.
Worked example 3: probability between two values
Section titled “Worked example 3: probability between two values”The heights cm of plants are modelled by
Find .
Use lower bound , upper bound , mean and standard deviation :
The interval extends standard deviations below the mean and standard deviations above it, so an answer greater than is reasonable.
Worked example 4: an upper-tail probability
Section titled “Worked example 4: an upper-tail probability”Scores on a test are modelled by
Find the probability that a randomly selected score exceeds .
The complement is often easiest:
Therefore
Using symmetry
Section titled “Using symmetry”The curve is symmetric about , so
Also,
For example, if , then
Self-check 2
Section titled “Self-check 2”Let . Find:
- ;
- ;
- .
Answer
The values and are one standard deviation below and above the mean.
Standardising with a z-score
Section titled “Standardising with a z-score”Any normal random variable can be transformed to the standard normal distribution
using
For an observed value , its z-score is
It measures signed distance from the mean in standard deviations:
- means the value equals the mean;
- means it is standard deviations above the mean;
- means it is standard deviations below the mean.
Standardising subtracts the location , then divides by the scale . The direction of an inequality is unchanged because .
Worked example 5: standardising a boundary
Section titled “Worked example 5: standardising a boundary”Let . Find by standardising.
Convert the boundary to a z-score:
Therefore
and hence
Modern calculators can obtain the same result directly from , but z-scores expose the reasoning and are essential in parameter problems.
Worked example 6: comparing different scales
Section titled “Worked example 6: comparing different scales”Amira scores on a test with mean and standard deviation . Ben scores on a different test with mean and standard deviation . Who performed better relative to their group?
Calculate both z-scores:
Since , Amira’s score lies further above her group’s mean. Therefore
This comparison uses the normal models, not the raw scores alone.
Self-check 3
Section titled “Self-check 3”For :
- find the z-score corresponding to ;
- use symmetry to find from .
Answer
The values and are equally far from the mean, so
to four decimal places.
Inverse normal problems
Section titled “Inverse normal problems”An inverse normal calculation starts with an area and finds the corresponding boundary. Most calculators expect the cumulative area to the left of the boundary.
If
enter left-tail area . If
enter left-tail area .
Worked example 7: finding a percentile
Section titled “Worked example 7: finding a percentile”The journey time minutes is modelled by
Find the time below which of journeys fall.
We need such that
Using inverse normal with area , mean and standard deviation gives
Therefore
Worked example 8: an upper-tail cutoff
Section titled “Worked example 8: an upper-tail cutoff”The diameter mm of a manufactured part is modelled by
The largest are rejected. Find the minimum diameter that is rejected.
If the cutoff is , then
An inverse normal function uses the area to the left:
Using area , mean and standard deviation gives
so the cutoff is
Worked example 9: a central interval
Section titled “Worked example 9: a central interval”Let . Find such that
The omitted probability is
Symmetry puts in each tail. The upper boundary therefore has cumulative area
The standard normal quantile is , so
and
Self-check 4
Section titled “Self-check 4”Weights kg are modelled by . Find:
- the 25th percentile;
- the cutoff exceeded by the heaviest .
Answer
For the 25th percentile, use left-tail area :
The heaviest lie above the cutoff, so its left-tail area is :
Finding unknown parameters
Section titled “Finding unknown parameters”Probability statements can determine or . Translate each statement into a standard normal quantile, then use
Do not round the quantile early. Retain several decimal places until the final answer.
Worked example 10: finding the mean
Section titled “Worked example 10: finding the mean”Suppose and
Find .
The standard normal value with cumulative probability is
Therefore
so
Hence
The result is sensible: has of the distribution below it, so it must lie above the mean.
Worked example 11: finding both parameters
Section titled “Worked example 11: finding both parameters”Let , where
and
Find and .
The corresponding z-values are
so
and
The probabilities are symmetric, so and are equally far from the mean:
Then
giving
Therefore
Worked example 12: a non-symmetric pair of conditions
Section titled “Worked example 12: a non-symmetric pair of conditions”Suppose and
Find and .
From inverse normal,
Hence
and
Subtracting the first equation from the second gives
Thus
and, substituting into the second equation,
Therefore
Self-check 5
Section titled “Self-check 5”Let and . Find .
Answer
The area below is , whose z-value is . Therefore
and
When is a normal model appropriate?
Section titled “When is a normal model appropriate?”A normal model may be reasonable when:
- the variable is continuous;
- the data are approximately symmetric and unimodal;
- frequencies taper smoothly on both sides of the centre;
- there are no strong outliers or physical boundaries close to the bulk of the data;
- contextual knowledge supports a stable population with many small sources of variation.
A histogram or other data presentation can reveal skewness, multiple peaks, gaps and outliers. Context still matters. A small sample can look roughly bell-shaped by chance, and some measurements cannot sensibly take the negative values that a normal model technically permits.
Normal approximation to the binomial
Section titled “Normal approximation to the binomial”If
then
When is large and is not too close to or , the binomial distribution is approximately normal:
A common practical condition is
although the exact convention can vary. Both expected successes and expected failures must be sufficiently numerous. Check the convention required by your course.
Why continuity correction is necessary
Section titled “Why continuity correction is necessary”The binomial variable takes integer values, whereas the normal variable takes all real values. To make their probability regions align, each integer is represented by an interval of width centred on it.
For example, is represented by
This adjustment is the continuity correction.
| Binomial event | Corrected normal event |
|---|---|
The strict or inclusive signs do not affect continuous probabilities. The boundary shift does.
Worked example 13: a cumulative binomial probability
Section titled “Worked example 13: a cumulative binomial probability”Let
Use a normal approximation to estimate .
First check suitability:
The approximating distribution has
and
Apply the continuity correction:
where . Therefore
Worked example 14: an interval
Section titled “Worked example 14: an interval”Each of seeds germinates independently with probability . Let be the number that germinate. Estimate
using a normal approximation.
Here
The suitability conditions hold because
Use
since
Continuity correction gives
Using normal CDF with standard deviation ,
so
Worked example 15: exactly one value
Section titled “Worked example 15: exactly one value”For , estimate using a normal approximation.
Both suitability values exceed :
Use
The single discrete value corresponds to the continuous interval from to :
Without continuity correction, the normal probability of exactly would be zero.
Self-check 6
Section titled “Self-check 6”Let . Use a normal approximation to estimate .
Answer
The conditions hold because and . Use
Since means , continuity correction gives
Therefore
Common misconceptions
Section titled “Common misconceptions”- Treating as . In , take the square root of the second parameter before using calculator fields that ask for standard deviation.
- Using curve height as probability. Probability is area under the curve over an interval.
- Thinking is positive. It is for a continuous random variable.
- Entering an upper-tail area into inverse normal. Convert it to the left-tail area first unless the calculator explicitly accepts a tail choice.
- Splitting a central area incorrectly. Divide the excluded area equally between the two tails.
- Assuming every symmetric data set is normal. A normal model must also be unimodal, smoothly tapering and contextually reasonable.
- Approximating a binomial without checking conditions. Verify both and .
- Forgetting continuity correction. Translate the integer event before using the continuous model.
Mixed self-check
Section titled “Mixed self-check”- . Find .
- For the same distribution, find if .
- and . Find .
- . State an approximating normal distribution and the corrected event for .
Answers
1. Using normal CDF,
2. The area to the left of is . Thus
3. Since ,
so
4. The mean and variance are
Therefore
is suitable, and
Exam checklist
Section titled “Exam checklist”Before finishing a normal distribution problem, check that you have:
- identified , and correctly;
- translated the words into a probability region;
- used a complement or symmetry correctly where needed;
- converted an upper-tail area before inverse normal;
- checked whether the numerical answer is plausible;
- checked approximation conditions and applied continuity correction for a binomial variable;
- stated probabilities between and and rounded only at the end.
Next steps
Section titled “Next steps”- Review choosing a distribution when deciding between statistical models.
- Apply normal probabilities to significance testing in normal hypothesis tests.
- Revisit hypothesis testing language before interpreting critical regions and p-values.
- Strengthen modelling judgements with probability modelling and outliers and cleaning data.