Skip to content

Probability modelling: assumptions, simulation and validation

A probability model is a mathematical description of a random situation. It specifies possible outcomes or events and assigns probabilities to them. We use the model to calculate predictions, then compare those predictions with evidence.

A model is not the real process. It is a deliberately simplified representation:

real situationassumptionsprobability modelprediction\boxed{\text{real situation}\longrightarrow\text{assumptions}\longrightarrow\text{probability model}\longrightarrow\text{prediction}}

Good modelling therefore needs more than correct arithmetic. You must decide whether the assumptions are reasonable, interpret results in context, and recognise when evidence suggests that the model should be revised.

You should be able to:

  • use fractions, decimals and percentages;
  • identify complements, unions and intersections using probability rules;
  • multiply probabilities for independent events and add probabilities for mutually exclusive events;
  • calculate a mean and interpret frequencies from data.

A useful modelling cycle has five stages.

  1. Define the process. State what is observed and which event matters.
  2. Make assumptions. For example, outcomes may be treated as equally likely or repeated trials as independent.
  3. Construct and use the model. Assign probabilities and calculate the required prediction.
  4. Compare with data. Decide whether discrepancies are plausible random variation or evidence of a weakness.
  5. Refine or report. Change an assumption if necessary, or state the conclusion with its limitations.

The cycle matters because several different models can describe the same real situation. The best choice depends on purpose and on the available evidence.

A cafe records whether its next customer buys tea, coffee or another drink. From recent data, the manager proposes

P(T)=0.35,P(C)=0.50.\operatorname{P}(T)=0.35, \qquad \operatorname{P}(C)=0.50.

The three categories are mutually exclusive and exhaustive, so their probabilities must sum to 11. Therefore

P(other)=1P(T)P(C)=10.350.50=0.15.\begin{aligned} \operatorname{P}(\text{other}) &=1-\operatorname{P}(T)-\operatorname{P}(C)\\ &=1-0.35-0.50\\ &=\boxed{0.15}. \end{aligned}

This model assumes that the recent data remain relevant. It may be poor on an unusually hot day, after a price change, or at a different time of day. The calculation can be flawless while the prediction is still unsuitable.

Probabilities can come from theoretical reasoning, observed data, expert judgement, or a combination of these.

If a finite sample space contains nn equally likely outcomes and event AA contains aa of them, then

P(A)=an.\boxed{\operatorname{P}(A)=\frac{a}{n}.}

The phrase equally likely is an assumption, not an automatic fact.

Worked example 2: when counting is justified

Section titled “Worked example 2: when counting is justified”

A fair six sided die is rolled. Find the probability of obtaining a prime number.

The sample space is

S={1,2,3,4,5,6}.S=\{1,2,3,4,5,6\}.

The prime outcomes are 2,3,52,3,5. Fairness makes all six faces equally likely, so

P(prime)=36=12.\operatorname{P}(\text{prime})=\frac{3}{6}=\boxed{\frac12}.

If the die were not known to be fair, merely counting three favourable faces would not justify the answer 1/21/2.

If event AA occurs rr times in nn trials, its relative frequency is

p^=rn.\boxed{\widehat{p}=\frac{r}{n}.}

This observed value can be used as an estimate of the unknown probability pp. The hat in p^\widehat p indicates an estimate, not an exact population value.

A drawing pin lands point upwards 138138 times in 200200 throws. Estimate the probability that it lands point upwards and predict the number of such results in the next 750750 throws.

The estimated probability is

p^=138200=0.69.\widehat p=\frac{138}{200}=0.69.

Using this estimate for future throws gives

750(0.69)=517.5.750(0.69)=\boxed{517.5}.

The number of actual successes must be an integer, but 517.5517.5 is an expected frequency, not a promised outcome. It would be reasonable to report about 518518 point upwards results.

The prediction assumes that throwing conditions and the drawing pin do not change and that the 200200 recorded throws are representative.

A machine produced 1717 faulty items among 400400 inspected items. Estimate the fault probability and predict the number of faulty items among the next 24002400 if conditions remain unchanged.

Answer p^=17400=0.0425.\widehat p=\frac{17}{400}=0.0425.

The predicted frequency is

2400(0.0425)=102.2400(0.0425)=\boxed{102}.

This is a model based prediction. It does not say that exactly 102102 items will be faulty.

For repeated trials under stable conditions, relative frequency tends to become more stable as the number of trials increases. This is the law of large numbers:

Xnnpas n becomes large,\frac{X_n}{n}\longrightarrow p \quad\text{as }n\text{ becomes large},

where XnX_n is the number of successes in nn trials and pp is the success probability.

This does not mean that the relative frequency moves closer to pp after every trial. It can move away temporarily. Nor does it mean that a success becomes more likely because several failures have just occurred.

Worked example 4: interpreting a short run

Section titled “Worked example 4: interpreting a short run”

A fair coin gives 77 heads in 1010 tosses. Is this evidence that the model P(H)=0.5\operatorname{P}(H)=0.5 is false?

The observed relative frequency is

710=0.7,\frac7{10}=0.7,

which differs from 0.50.5. However, short sequences naturally vary. In fact, under the fair coin model,

P(exactly 7 heads)=(107)(12)7(12)3=12010240.117.\begin{aligned} \operatorname{P}(\text{exactly }7\text{ heads}) &=\binom{10}{7}\left(\frac12\right)^7\left(\frac12\right)^3\\ &=\frac{120}{1024}\\ &\approx0.117. \end{aligned}

This particular result is not especially rare. Ten tosses provide weak evidence about fairness. More observations would make the estimate more informative.

If an event has modelled probability pp in each of nn trials, its expected frequency is

np.\boxed{np.}

It is the long run average count over many repetitions of the whole set of nn trials. It need not be an integer and need not occur in any one set of trials.

Worked example 5: an expectation that is not possible

Section titled “Worked example 5: an expectation that is not possible”

For a fair die rolled 2020 times, the expected number of sixes is

20(16)=1033.33.20\left(\frac16\right)=\frac{10}{3}\approx3.33.

It is impossible to observe 3.333.33 sixes. The value means that if the experiment of 2020 rolls were repeated many times, the mean number of sixes per experiment would approach 10/310/3.

A seed has probability 0.820.82 of germinating. Give the expected number that germinate from 150150 seeds. Explain why the actual number need not equal your answer.

Answer 150(0.82)=123.150(0.82)=\boxed{123}.

The probability model describes long run behaviour. Random variation means one batch can produce more or fewer than 123123 germinations.

An assumption should be specific enough to test or criticise. Saying only that a model is unrealistic earns little credit.

AssumptionMathematical consequenceA possible failure
Outcomes are equally likelyProbability can be found by countingA spinner has unequal sectors or an off-centre pivot
Trials are independentJoint probabilities can be multipliedWeather on consecutive days is related
Success probability is constantOne value of pp applies throughoutA player’s fatigue changes their scoring chance
Categories are exhaustiveTheir probabilities sum to 11An overlooked response category exists
Data are representativeRelative frequency estimates the target probabilityData come from one untypical time or group
Conditions remain stablePast data can predict future behaviourA process, population or policy changes

Independence and constant probability are different assumptions. A machine can retain a constant long run fault rate while faults occur in clusters, which violates independence.

Worked example 6: criticising an independence assumption

Section titled “Worked example 6: criticising an independence assumption”

A commuter is late on 20%20\% of working days. A simple model treats lateness on different days as independent. Find the modelled probability that the commuter is late on both Monday and Tuesday, then assess the model.

Under the stated model,

P(LMLT)=P(LM)P(LT)=(0.2)(0.2)=0.04.\operatorname{P}(L_M\cap L_T) =\operatorname{P}(L_M)\operatorname{P}(L_T) =(0.2)(0.2) =\boxed{0.04}.

The independence assumption may be doubtful. Severe weather, engineering work, illness, or a disrupted rail service can affect both days. If lateness clusters, the true probability of two late days could exceed 0.040.04.

The useful criticism identifies the assumption, gives a contextual reason it may fail, and states how the prediction could be affected.

Worked example 7: sampling without replacement

Section titled “Worked example 7: sampling without replacement”

A batch contains 10001000 components, of which 4040 are faulty. Two are selected without replacement. A quick model treats the selections as independent with fault probability 0.040.04.

The approximate probability that both are faulty is

(0.04)2=0.0016.(0.04)^2=0.0016.

The exact probability is

401000×399990.0015616.\frac{40}{1000}\times\frac{39}{999} \approx0.0015616.

The selections are not exactly independent because the first selection changes the batch. The approximation is nevertheless close because the sample of 22 is tiny relative to the population of 10001000.

This illustrates an important principle: an assumption can be false in a literal sense but still give a useful approximation.

A simulation imitates a random process using random numbers. It is valuable when exact calculation is difficult, when a model has several stages, or when we want to study its long run behaviour.

A valid simulation must include:

  • a clear mapping from random numbers to outcomes;
  • probabilities that match the proposed model;
  • one complete definition of a trial;
  • a large number of repetitions;
  • a recorded quantity that answers the question.

If an event has probability 0.370.37, two digit random integers from 0000 to 9999 could represent it by assigning 0000 to 3636 to the event and 3737 to 9999 to its complement. This uses exactly 3737 of the 100100 equally likely values.

A basketball player scores each free throw with probability 0.720.72. Design a simulation to estimate the probability that the player scores at least 44 of 55 throws, assuming independence.

Use two digit random integers from 0000 to 9999.

  1. Let 0000 to 7171 represent a score and 7272 to 9999 represent a miss.
  2. Generate five integers. These represent one set of five throws.
  3. Record a success for the set if at least four integers represent scores.
  4. Repeat the set many times, say 1000010\,000 times.
  5. Estimate the required probability by
number of successful sets10000.\frac{\text{number of successful sets}}{10\,000}.

The simulation itself does not establish that 0.720.72 is the correct probability or that throws are independent. It explores the consequences if those assumptions hold.

For comparison, the exact probability under this model is

P(X4)=(54)(0.72)4(0.28)+(55)(0.72)50.577.\begin{aligned} \operatorname{P}(X\geq4) &=\binom54(0.72)^4(0.28)+\binom55(0.72)^5\\ &\approx\boxed{0.577}. \end{aligned}

Simulation estimates fluctuate around this value. More repetitions usually reduce, but do not eliminate, simulation error.

A customer buys a warranty with probability 0.130.13. Describe a simulation using random digits 00 to 99 to estimate the probability that at least one of the next four customers buys one.

Answer

One digit cannot represent probability 0.130.13 exactly, so combine digits in pairs to form equally likely integers 0000 to 9999.

  • Assign 0000 to 1212 to a warranty purchase and 1313 to 9999 to no purchase.
  • Generate four pairs of digits for one trial.
  • Record whether at least one pair is from 0000 to 1212.
  • Repeat many times.
  • Divide the number of recorded successes by the number of trials.

The method assumes customers make independent decisions with constant purchase probability 0.130.13.

Suppose a model assigns probabilities p1,p2,,pkp_1,p_2,\ldots,p_k to kk categories. In nn observations, the expected frequency in category ii is

Ei=npi.E_i=np_i.

Compare expected and observed frequencies. Some difference is inevitable because of random variation. A close match supports the model’s usefulness but does not prove it true. A large or systematic discrepancy prompts questions about assumptions, data quality, or changing conditions.

A die is rolled 120120 times, giving the following results.

Score112233445566
Observed frequency141418181717202023232828

Under a fair die model, every expected frequency is

120(16)=20.120\left(\frac16\right)=20.

The differences, observed minus expected, are

6, 2, 3, 0, 3, 8.-6,\ -2,\ -3,\ 0,\ 3,\ 8.

The six appears more often and the one less often than expected, but discrepancies alone do not prove bias. A careful conclusion is:

The data show some departure from the fair die model, particularly for scores 11 and 66. This could be random variation, so more data or a formal goodness of fit procedure would be needed before concluding that the die is biased.

At A level, the key skill here is measured interpretation. Do not say that observed and expected frequencies should be identical.

Model choice depends on the variable and mechanism, not only on the appearance of data.

A bakery wants to model the number of loaves sold each weekday. It considers:

  • Model A: every weekday has the same demand distribution;
  • Model B: Mondays, Fridays and other weekdays have separate distributions.

Model A is simpler and needs fewer data. Model B can represent a genuine day effect but needs enough observations for each category. If Friday demand is consistently higher, Model A may systematically underpredict Fridays and overpredict quieter days.

A sensible process is to fit both models using earlier data, test their predictions on later data, and prefer the simpler model unless the added complexity produces a worthwhile improvement.

”There are two outcomes, so each has probability 1/21/2

Section titled “”There are two outcomes, so each has probability 1/21/21/2””

False unless the two outcomes are equally likely. Winning the lottery and not winning are two outcomes with very different probabilities.

”The expected frequency is what will happen”

Section titled “”The expected frequency is what will happen””

Expected frequency is a long run mean. Actual frequencies vary randomly.

”A larger sample removes all uncertainty”

Section titled “”A larger sample removes all uncertainty””

A larger representative sample usually reduces random sampling error. It does not repair bias, poor measurement, dependence, or changing conditions.

For independent fair coin tosses,

P(HTTTTT)=12.\operatorname{P}(H\mid TTTTT)=\frac12.

Past tosses do not compensate for an imbalance. Believing that they must is the gambler’s fallacy.

Different models can make similar predictions, especially with limited data. Agreement means the model may be adequate for a stated purpose, not that every assumption is literally true.

A company claims that 90%90\% of parcels arrive the next day. In a sample of 8080 parcels, 6868 arrive the next day.

a. Find the observed relative frequency.

b. Find the expected number of next day arrivals under the company’s model.

c. State two assumptions needed when applying the model to future parcels.

d. Explain why the sample does not by itself prove the claim false.

a.

p^=6880=0.85.\widehat p=\frac{68}{80}=\boxed{0.85}.

b.

80(0.9)=72.80(0.9)=\boxed{72}.

c. Suitable assumptions include:

  • the sampled parcels are representative of the parcels covered by the claim;
  • delivery conditions remain stable;
  • one parcel’s arrival does not materially affect another’s;
  • the meaning and recording of “next day” are consistent.

d. The observed count is 44 below the expected count, but samples vary randomly. A formal test would measure how unusual 6868 or fewer is under the p=0.9p=0.9 model. This leads to binomial hypothesis testing.

A spinner is claimed to land on red with probability 0.40.4. It lands on red 9393 times in 250250 spins.

  1. Find the experimental probability of red.
  2. Find the expected red frequency under the proposed model.
  3. State one assumption behind using the model for another spinner.
  4. Explain briefly whether the result disproves the model.
Answer
  1. The experimental probability is
93250=0.372.\frac{93}{250}=\boxed{0.372}.
  1. The expected frequency is
250(0.4)=100.250(0.4)=\boxed{100}.
  1. For example, the other spinner must be manufactured and operated under sufficiently similar conditions. Alternatively, the spins should be independent and the probability should remain constant.

  2. No. The observed frequency differs from the expectation by 77, but random variation is expected. The result is evidence to assess, not automatic disproof. More data or a formal probability calculation would be needed for a stronger conclusion.

You should now be able to:

  • distinguish a probability model from the real process;
  • assign probabilities from equally likely outcomes or relative frequencies;
  • calculate and interpret expected frequencies;
  • explain long run relative frequency without invoking the gambler’s fallacy;
  • identify independence, constant probability, representativeness and stability assumptions;
  • design a random number simulation with a correct outcome mapping;
  • compare observed and expected results without overclaiming;
  • state contextual limitations and suggest sensible refinements.