Skip to content

Mathematical modelling

A mathematical model represents a real situation using variables, equations, graphs, diagrams or probability distributions. It deliberately ignores some detail so that mathematics can answer a useful question.

For example, the equation

d=vtd=vt

models distance dd travelled at constant speed vv for time tt. It is not universally true. It is useful only when the speed is constant, the units are consistent, and dd means distance along the route.

The central habit of modelling is therefore not merely to calculate. It is to ask:

  1. What do the symbols mean?
  2. Which assumptions make the mathematics possible?
  3. Is the result sensible in the original situation?

You should be able to:

  • substitute into and rearrange formulae;
  • read tables and graphs;
  • round numbers appropriately;
  • convert units and compound measures;
  • distinguish exact values from approximations.

A complete modelling solution moves between the real situation and mathematics.

real problemassumptions and variablesmathematical modelsolutioninterpretation and evaluation\boxed{ \text{real problem} \longrightarrow \text{assumptions and variables} \longrightarrow \text{mathematical model} \longrightarrow \text{solution} \longrightarrow \text{interpretation and evaluation} }

If the answer is not accurate enough for its purpose, refine the assumptions or choose a different model and repeat the cycle.

StageQuestions to ask
DefineWhat is known? What must be found? What are the units and allowed values?
SimplifyWhich effects are important? Which can reasonably be ignored?
FormulateWhich variables and relationships represent the situation?
SolveWhich algebraic, graphical, numerical or statistical method applies?
InterpretWhat does the mathematical answer mean in context?
ValidateDoes it fit observations, units, constraints and common sense?
RefineWhich assumption or parameter should change?

In an examination, these stages may be compressed into one question. Your written solution should still make them visible.

A variable needs a meaning, a unit and often a domain. Compare

x=5x=5

with

t=5 minutes after the tank begins draining.t=5\text{ minutes after the tank begins draining}.

The second statement can be interpreted and checked. The first cannot.

Suppose nn is the number of complete rows of seats. Then

n{0,1,2,},n\in\{0,1,2,\ldots\},

not merely n0n\geq0. A model may produce n=14.7n=14.7, but the context permits only a whole number. Whether to use 1414 or 1515 depends on the question, not on the usual rounding rule.

Worked example 1: identify the variables and domain

Section titled “Worked example 1: identify the variables and domain”

A taxi charges a fixed fee of £3.20\pounds 3.20 plus £1.80\pounds 1.80 per mile. Model the fare for a journey of mm miles.

Let

m=journey length in miles,C=fare in pounds.m=\text{journey length in miles}, \qquad C=\text{fare in pounds}.

The model is

C=3.20+1.80m,m0.\boxed{C=3.20+1.80m}, \qquad m\geq0.

Here 3.203.20 is the fare at m=0m=0, and 1.801.80 is the rate of change of fare with distance, measured in pounds per mile.

If the meter records each completed tenth of a mile, the actual domain is discrete:

m{0,0.1,0.2,}.m\in\{0,0.1,0.2,\ldots\}.

That operational detail changes the model even though the algebraic formula looks the same.

An assumption is a statement accepted while building the model. Assumptions reduce complexity, but each one limits when the result can be trusted.

Common assumptions include:

  • an object is a particle, so its size and rotation are ignored;
  • a string is light and inextensible;
  • air resistance is negligible;
  • acceleration is constant;
  • observations are independent;
  • a sample is representative of a population;
  • a percentage rate remains constant;
  • a relationship seen within the data continues beyond it.

Good assumptions are specific. Write air resistance is negligible, not conditions are normal.

Worked example 2: expose the hidden assumptions

Section titled “Worked example 2: expose the hidden assumptions”

A runner completes 400400 m in 5050 s. Using

v=dt,v=\frac{d}{t},

we find

v=40050=8 m s1.v=\frac{400}{50}=8\text{ m s}^{-1}.

This is the runner’s mean speed, not necessarily the speed at every instant. To use d=vtd=vt to predict that the runner covers 640640 m in the next 8080 s, we must assume that the same mean speed continues:

d=8×80=640 m.d=8\times80=640\text{ m}.

The prediction may fail because of fatigue, acceleration from rest, bends or changing conditions. The arithmetic can be correct while the model is poor.

Different patterns suggest different mathematical structures.

BehaviourTypical modelRecognising feature
constant amount added per unity=a+bxy=a+bxconstant gradient
constant factor per intervaly=Abxy=Ab^xconstant percentage change
quantity proportional to a powery=Axny=Ax^nscaling xx scales yy predictably
periodic variationy=a+bsin(cx+d)y=a+b\sin(cx+d)repeating cycle
constant accelerationv=u+atv=u+atvelocity changes linearly with time
random outcomeprobability distributionprobabilities describe uncertainty

This table is a starting point, not proof. Several models may fit a small data set. Context and validation decide which is useful.

Two savings plans both start with £1000\pounds 1000.

  • Plan A adds £60\pounds 60 each year.
  • Plan B adds 6%6\% of the current balance each year.

After nn years, Plan A is linear:

An=1000+60n.A_n=1000+60n.

Plan B is exponential because each balance is multiplied by 1.061.06:

Bn=1000(1.06)n.B_n=1000(1.06)^n.

After 1010 years,

A10=1000+60(10)=1600,A_{10}=1000+60(10)=1600,

whereas

B10=1000(1.06)10=1790.847.B_{10}=1000(1.06)^{10}=1790.847\ldots.

Thus the models predict £1600\pounds1600 and £1790.85\pounds1790.85 respectively.

The phrase adds $6\%$ does not mean add 6060 every year. The percentage acts on a changing balance.

Every term added in an equation must have the same units. In

s=ut+12at2,s=ut+\frac12at^2,

if uu is in m s1\text{m s}^{-1}, aa in m s2\text{m s}^{-2} and tt in seconds, then

[ut]=(m s1)(s)=m[ut]=(\text{m s}^{-1})(\text{s})=\text{m}

and

[at2]=(m s2)(s2)=m.[at^2]=(\text{m s}^{-2})(\text{s}^2)=\text{m}.

Both terms have units of length, as ss must.

Unit analysis can reject an impossible formula, but it cannot prove a formula correct. Both s=ut+at2s=ut+at^2 and s=ut+12at2s=ut+\frac12at^2 are dimensionally consistent.

Worked example 4: repair inconsistent units

Section titled “Worked example 4: repair inconsistent units”

A cyclist travels at 18 km h118\text{ km h}^{-1} for 4040 seconds. Find the distance travelled under a constant speed assumption.

Convert the speed first:

18 km h1=18×10003600 m s1=5 m s1.18\text{ km h}^{-1} =18\times\frac{1000}{3600}\text{ m s}^{-1} =5\text{ m s}^{-1}.

Then

d=vt=5×40=200 m.d=vt=5\times40=200\text{ m}.

Therefore the model predicts

d=200 m.\boxed{d=200\text{ m}}.

Multiplying 1818 directly by 4040 combines hours with seconds and has no valid interpretation.

A mathematical solution is not automatically a contextual answer. Check:

  • sign: can the quantity be negative?
  • domain: must it be a whole number or lie in a fixed interval?
  • scale: is the order of magnitude plausible?
  • precision: do the data justify the stated accuracy?
  • meaning: which root or solution answers the question?

Worked example 5: reject an inadmissible solution

Section titled “Worked example 5: reject an inadmissible solution”

The height of a ball is modelled by

h=1.2+14t4.9t2,h=1.2+14t-4.9t^2,

where hh is in metres and tt is seconds after release. Find when the model says the ball reaches the ground.

Set h=0h=0:

1.2+14t4.9t2=0,1.2+14t-4.9t^2=0,

or

4.9t214t1.2=0.4.9t^2-14t-1.2=0.

Using the quadratic formula,

t=14±(14)24(4.9)(1.2)2(4.9)=14±219.529.8.t=\frac{14\pm\sqrt{(-14)^2-4(4.9)(-1.2)}}{2(4.9)} =\frac{14\pm\sqrt{219.52}}{9.8}.

This gives

t2.940ort0.0833.t\approx2.940 \qquad\text{or}\qquad t\approx-0.0833.

The model begins at t=0t=0, so the negative value describes an extrapolation before release and is outside the contextual domain. Hence

t2.94 s.\boxed{t\approx2.94\text{ s}}.

Do not call the negative root wrong. It solves the equation, but it is inadmissible for this question.

Worked example 6: rounding depends on purpose

Section titled “Worked example 6: rounding depends on purpose”

A minibus holds 1616 passengers. How many minibuses are required for 9393 passengers?

9316=5.8125.\frac{93}{16}=5.8125.

Five minibuses are insufficient, so the answer must be rounded up:

6 minibuses.\boxed{6\text{ minibuses}}.

By contrast, if 9393 identical items are packed into complete boxes of 1616, the number of full boxes is

5,\boxed{5},

with 1313 items left. The same calculation leads to different rounding because the questions ask for different things.

Validation compares the model with reality or with information not used to construct it. Useful checks include:

  1. substitute a known case;
  2. compare predictions with fresh observations;
  3. inspect residuals, the differences observedpredicted\text{observed}-\text{predicted};
  4. test boundary and extreme cases;
  5. check whether parameters have sensible signs and sizes;
  6. reconsider the assumptions.

Water depth dd in a tank is modelled by

d=1.800.12t,d=1.80-0.12t,

where dd is in metres and tt is in minutes. At t=5t=5, the observed depth is 1.241.24 m.

The predicted depth is

d=1.800.12(5)=1.20 m.d=1.80-0.12(5)=1.20\text{ m}.

The residual is

observedpredicted=1.241.20=0.04 m.\text{observed}-\text{predicted}=1.24-1.20=0.04\text{ m}.

So the model underestimates the observed depth by 0.040.04 m.

Whether this is acceptable depends on purpose. An error of 44 cm may be harmless for a rough display and unacceptable for an overflow warning system.

The model also predicts d=0d=0 when

1.800.12t=0,1.80-0.12t=0,

so t=15t=15 minutes. For t>15t>15, it predicts negative depth, which is physically impossible. Its sensible domain is at most

0t15.0\leq t\leq15.

Interpolation predicts within the range of observed data. Extrapolation predicts outside that range.

If temperatures have been recorded for 10t3010\leq t\leq30 minutes, estimating at t=18t=18 is interpolation, while estimating at t=80t=80 is extrapolation.

Extrapolation is usually less reliable because:

  • the relationship may change;
  • a physical limit may be reached;
  • omitted effects may become important;
  • small parameter errors can produce large long term errors.

For example, a linear population model with positive gradient eventually predicts unlimited growth. Resource limits make that implausible over a sufficiently long period.

A model is sensitive to a parameter if a small change in that parameter produces a large change in the prediction.

A population after 2020 years is modelled by

P=5000(1+r)20,P=5000(1+r)^{20},

where rr is the annual growth rate as a decimal.

For r=0.030r=0.030,

P=5000(1.03)20=9030.56.P=5000(1.03)^{20}=9030.56\ldots.

For the slightly larger estimate r=0.032r=0.032,

P=5000(1.032)20=9386.88.P=5000(1.032)^{20}=9386.88\ldots.

A change of only 0.20.2 percentage points in the annual rate changes the twenty year prediction by about

9386.889030.56=356.32.9386.88-9030.56=356.32.

This does not make the model useless. It means the rate must be estimated carefully and long term predictions should be reported with appropriate caution.

Possible refinements include:

  • allowing a parameter to vary with time;
  • replacing a straight line with a curve;
  • splitting the domain into different regimes;
  • adding an omitted force or constraint;
  • collecting more representative data;
  • stating a narrower domain of validity.

Each refinement adds complexity. Prefer the simplest model accurate enough for the intended use.

The model gave a number, so the answer is reliable

Section titled “The model gave a number, so the answer is reliable”

Calculation only shows what follows from the model. Reliability also depends on assumptions, data quality, domain and purpose.

More decimal places make a prediction more accurate

Section titled “More decimal places make a prediction more accurate”

Decimal places show numerical precision, not model accuracy. If a length was measured as 2.42.4 m, reporting a prediction as 17.63829117.638291 m usually gives unjustified precision.

An association can support prediction without establishing causation. A third variable or coincidence may explain the pattern. See correlation and regression.

All solutions of the equation answer the problem

Section titled “All solutions of the equation answer the problem”

An equation may have negative, non integer or out of range solutions. Apply the contextual domain after solving.

An unrealistic assumption makes the whole model worthless

Section titled “An unrealistic assumption makes the whole model worthless”

All models simplify. The proper question is whether the assumption causes unacceptable error for the required purpose.

For an extended modelling question, use this compact structure:

  1. Define: Let $t$ be time in seconds and $h$ be height in metres.
  2. Assume: Assume constant acceleration and negligible air resistance.
  3. Model: write the equation and its domain.
  4. Solve: show the mathematics with units.
  5. Interpret: select admissible solutions and answer in context.
  6. Evaluate: state a limitation and its likely effect.

A useful evaluation connects cause to consequence. For example, ignoring air resistance makes the predicted range too large because drag reduces the horizontal speed.

Merely writing the model is unrealistic gives no mathematical insight.

A tap fills a container at a constant rate of 0.350.35 litres per second. The container initially holds 2.02.0 litres.

  1. Write a model for the volume VV after tt seconds.
  2. State the units of the gradient.
  3. Find the predicted volume after 1212 seconds.
Answer

Let VV be volume in litres and tt be time in seconds. Then

V=2.0+0.35t,t0,V=2.0+0.35t, \qquad t\geq0,

until the container becomes full or the flow changes. The gradient has units litres per second. At t=12t=12,

V=2.0+0.35(12)=6.2 litres.V=2.0+0.35(12)=6.2\text{ litres}.

A model for the number of crates required gives n=23.12n=23.12. Each crate holds at most the stated capacity. What value of nn should be used?

Answer

Use

n=24.n=24.

The number of crates must be a whole number, and 2323 crates would not provide enough capacity.

A linear model fitted to the height of a plant during its first six weeks predicts a height of 9.49.4 m after five years. Give two reasons to doubt this prediction.

Answer

Five years lies far outside the observed six week range, so this is substantial extrapolation. The plant’s growth rate is unlikely to remain constant because maturity, seasonal conditions, disease and limited resources can change it. The linear relationship should therefore not be assumed to continue for five years.

A quantity starts at 8080. Model L adds 1212 each period. Model E increases by 12%12\% each period. Find both predictions after 55 periods.

Answer

For the linear model,

L=80+12(5)=140.L=80+12(5)=140.

For the exponential model,

E=80(1.12)5=140.987141.E=80(1.12)^5=140.987\ldots\approx141.

The similar five period answers do not make the models equivalent. Their predictions separate increasingly over time.

A model predicts the stopping distance ss of a car from s=0.08v2s=0.08v^2, where ss is in metres and vv is in metres per second. A student substitutes v=72v=72 for a speed of 72 km h172\text{ km h}^{-1}. Explain and correct the error.

Answer

The formula requires metres per second, but 7272 is in kilometres per hour. Convert:

72 km h1=72×10003600=20 m s1.72\text{ km h}^{-1}=72\times\frac{1000}{3600}=20\text{ m s}^{-1}.

Therefore

s=0.08(20)2=32 m.s=0.08(20)^2=32\text{ m}.

The transferable principle is simple: solve the mathematics, then judge the solution as a statement about the real situation.