Mathematics · IGCSE 0580 · §9.1–9.6

Statistics

A page of raw data says almost nothing until it is summarised — a few well-chosen numbers and the right diagram turn a list into an argument.

Mathematics · 0580 Extended Topic 12 of 12

Averages & Range

An average is a single number standing in for a whole data set. Three are used at IGCSE, each a different sense of "typical"; the range measures how spread out the values are.

Definition
Average
A single value chosen to represent a whole data set. Mean, median and mode are three different averages.
Definition
Range
Largest value − smallest value. A measure of spread, not an average.

The three averages and the range

The mean shares the total equally: mean = (sum of values) ÷ (number of values). The median is the middle value once the data is ordered. The mode occurs most often. The range is the gap between the extremes, highest − lowest.

Worked example: a team scores 1, 3, 0, 2, 3, 4, 2, 3, 1 goals across nine matches. Find the mean, median, mode and range. Step 1: order the data — 0, 1, 1, 2, 2, 3, 3, 3, 4. Step 2: mean = 19 ÷ 9 = 2.11 (3 s.f.). Step 3: median = 5th value = 2; mode = 3 (three times); range = 4 − 0 = 4 — each answers a different question. Mean 2.11, median 2, mode 3, range 4.

The mean from a frequency table

When values repeat, a frequency table is quicker. Multiply each value x by its frequency f, total the products, and divide by the total frequency.

Worked example: in a class of 40, the number of pets owned is 0 (×8), 1 (×14), 2 (×10), 3 (×6), 4 (×2). Find the mean. Step 1: Σf = 8 + 14 + 10 + 6 + 2 = 40. Step 2: Σfx = 0 + 14 + 20 + 18 + 8 = 60. Step 3: mean = 60 ÷ 40 = 1.5 pets per student.

Examiner note
For the median, order the data first, every time. The median is the value in position (n + 1) ÷ 2.
Why this matters
One very high salary drags the mean upward but leaves the median untouched — which is why pay is usually reported as a median.

Grouped Data & Spread

Once data is grouped into classes, the individual values are gone: we can only estimate the mean, and we measure spread with the interquartile range, which ignores the extremes.

Definition
Interquartile range
IQR = Q₃ − Q₁. The spread of the middle 50%, so extreme values do not distort it.

Estimating the mean of grouped data

We assume every value sits at its class midpoint, then work out estimated mean = Σfx ÷ Σf as before, where x is the midpoint of each class. The modal class is the interval with the greatest frequency.

Definition
Modal class
The class interval with the highest frequency — the grouped-data version of the mode.

Worked example: the mass of 60 apples is grouped as 100<m≤120 (f = 8), 120<m≤140 (f = 20), 140<m≤160 (f = 22), 160<m≤180 (f = 10). Estimate the mean and state the modal class. Step 1: midpoints 110, 130, 150, 170. Step 2: Σfx = 880 + 2600 + 3300 + 1700 = 8480. Step 3: estimated mean = 8480 ÷ 60 = 141.3 g (1 d.p.). Step 4: the highest frequency is 22, so the modal class is 140<m≤160.

Quartiles and the interquartile range

For a list, find the median, then Q₁ is the middle of the lower half and Q₃ the middle of the upper half. The IQR, Q₃ − Q₁, measures the spread of the central half.

Worked example: eleven ordered test scores are 4, 6, 7, 7, 8, 9, 11, 12, 12, 14, 15. Find the quartiles and the IQR. Step 1: median Q₂ = 6th value = 9. Step 2: lower half 4, 6, 7, 7, 8 → Q₁ = 7; upper half 11, 12, 12, 14, 15 → Q₃ = 12. Step 3: IQR = 12 − 7 = 5. Q₁ = 7, Q₃ = 12, IQR = 5.

Definition
Quartiles
The three values (Q₁, Q₂, Q₃) that cut ordered data into four equal parts. Q₂ is the median.
Examiner note
A grouped-data mean is only an estimate — the exact values are unknown. Always write "estimate" and use the midpoint of each class.

Charts & Diagrams

The right diagram makes a pattern obvious at a glance. Each has a job: bar charts compare categories, pie charts show proportions of a whole, and stem-and-leaf diagrams keep every value while revealing the shape of the data.

Definition
Pie chart
A circle split into sectors, each sector's angle proportional to the frequency it represents.
Definition
Stem-and-leaf
Displays every value while showing shape. Leaves must be ordered and a key is compulsory.

Pie charts: angles from frequencies

A full circle is 360° and represents the whole data set. Each category takes a slice whose angle is its share of the total: sector angle = (frequency ÷ total) × 360°, and the angles must sum to 360°. Composite (stacked) and dual (side-by-side) bar charts do the same comparison job for two sub-groups at once.

144° Football — 24 → 144° Netball — 15 → 90° Swimming — 12 → 72° Other — 9 → 54° 60 students · angles total 360°
FIG 9.1 A pie chart of favourite sports; each sector angle is the category's share of 360°.

Worked example: of 60 students, 24 chose football. Find the angle of the football sector. Step 1: fraction = 24 ÷ 60. Step 2: angle = (24 ÷ 60) × 360° = 144°. Football sector = 144°.

Examiner note
A stem-and-leaf diagram with no key, or with unordered leaves, loses a mark even when the numbers are right.
Why this matters
To compare two data sets, quote one average and one measure of spread — and interpret both in context, not just state them.

Scatter & Correlation

A scatter diagram plots two variables against each other, one point per item, to reveal whether they are linked. If the points trend along a line, a line of best fit turns that trend into predictions.

Definition
Line of best fit
A single ruled straight line drawn by eye, with roughly as many points above it as below.

Reading correlation

Points rising left to right show positive correlation; falling points show negative correlation; a shapeless cloud shows zero correlation. The tighter the points cluster to a line, the stronger the correlation. Plot points clearly, as small crosses.

Definition
Correlation
The way two variables change together: positive (both rise), negative (one rises as the other falls), or zero (no link).
revision hours test score line of best fit
FIG 9.2 Positive correlation: more revision hours tend to go with higher scores. The ruled line is drawn by eye.

Worked example: a student revised for 5 hours but was absent for the test. Use the line of best fit to estimate their score. Step 1: find 5 hours on the horizontal axis and read up to the line. Step 2: read across to the score axis — about 62 marks. Step 3: this is an estimate — 5 hours lies inside the data, so it is reasonable. Estimated score ≈ 62.

Examiner note
Correlation is not cause. A strong line does not prove one variable makes the other change — say so if asked.
Why this matters
Predicting far outside the plotted range (extrapolation) is unreliable — the pattern may not continue.

Cumulative Frequency

When data is grouped, the median and quartiles cannot be found exactly — but a cumulative frequency curve lets us estimate them by reading values straight off the graph. The same grouped table also builds a histogram, covered next.

Definition
Cumulative frequency
A running total of the frequencies — how many values fall at or below each class boundary.

Building and reading the curve

Form a running total of the frequencies, then plot each cumulative total against the upper boundary of its class and join the points with a smooth curve. To estimate the median of n values, go to n ÷ 2 on the cumulative axis, across to the curve and down to the value.

time (seconds) cumulative freq n ÷ 2 = 40 median
FIG 9.3 Cumulative frequency curve for 80 runners; reading across from 40 gives the median time.

Worked example: for the 80 runners above, estimate the median and the interquartile range from the curve. Step 1: median at 40 — read off ≈ 57 s. Step 2: Q₁ at 20 → 50 s; Q₃ at 60 → ≈ 65 s. Step 3: IQR = 65 − 50 = 15 s. Median ≈ 57 s, IQR ≈ 15 s.

Examiner note
Plot each point at the upper class boundary, and join with a smooth curve — never straight segments.
Examiner note
On a cumulative frequency curve read the median at n ÷ 2, Q₁ at n ÷ 4 and Q₃ at 3n ÷ 4.

Histograms

When class intervals have unequal widths, a bar chart of raw frequency misleads — a wide class looks bigger simply because it is wide. A histogram fixes this by making area, not height, represent frequency.

Frequency density

Divide each frequency by its class width to get the frequency density, and use that as the bar height: frequency density = frequency ÷ class width, so frequency = frequency density × class width. Now a bar's area equals its frequency, so classes of different widths are compared honestly.

Definition
Frequency density
Frequency ÷ class width. It is the height of each histogram bar, so that area represents frequency.
height (cm) frequency density 1.5 2.5 2.0 0.67
FIG 9.4 A histogram of 100 plant heights; unequal class widths, so bar height is frequency density.

Worked example: in the class 20<h≤40, the frequency is 40. Find the frequency density. Step 1: class width = 40 − 20 = 20. Step 2: frequency density = 40 ÷ 20 = 2.0. Step 3: a wide class with a modest density — area still records all 40 values.

Examiner note
Label the vertical axis "Frequency density", never "Frequency". Plotting raw frequency for unequal classes is the classic error.
Examiner note
Bars touch with no gaps — the horizontal axis is a continuous number line, not separate categories.

Exam advice

Common mistakes

Plotting raw frequency on a histogram with unequal classes
The height must be frequency density = frequency ÷ class width; using frequency loses the method and answer marks.
Reading the median off the frequency axis, not from n ÷ 2
On a cumulative frequency curve the median is at n ÷ 2 up the vertical axis, then across to the curve.
Calling a grouped-data mean exact
It is only an estimate from midpoints; presenting it as exact, or omitting the midpoint step, loses marks.
Stem-and-leaf with no key or unordered leaves
Both are required; a diagram missing either is not awarded full marks even if the data is correct.

Model answer

The table shows the mass, m grams, of 60 apples: 100<m≤120 (8), 120<m≤140 (20), 140<m≤160 (22), 160<m≤180 (10). Calculate an estimate of the mean mass.
[4 marks]
M1
Use the midpoint of each class
midpoints 110, 130, 150, 170
M1
Multiply by frequency and total
Σfx = 880 + 2600 + 3300 + 1700 = 8480
M1
Divide by the total frequency
8480 ÷ 60
A1
State the estimate to a sensible accuracy
= 141.3 g (1 d.p.), written as an estimate

Recall checklist

  • State the mean, median, mode and range of a list.
  • Calculate a mean from a frequency table.
  • Estimate the mean of grouped data using midpoints.
  • Identify the modal class and the quartiles.
  • Calculate the interquartile range, Q₃ − Q₁.
  • Find a pie-chart angle from a frequency.
  • Read the median and quartiles from a cumulative frequency curve.
  • Calculate frequency density and draw a histogram.

Every Mathematics topic, in one PDF you keep

Print it, write on it, revise with no wifi and no ads. One payment — not a subscription.

Get the Mathematics PDF

Ready to test this topic? Practise with Mathematics past papers and mark schemes →