The Sample Mean
Lesson 21 left you with a sample and an unknown population. Now you compute something from that sample. The sample mean $\bar{x}$ is the obvious thing to compute and the arithmetic takes one line, so the whole lesson turns on the second idea: draw a fresh sample the same size and you get a different $\bar{x}$ . Not because anyone made a mistake. That variation is not noise to be apologised for, it is the object the next three lessons study.
Write your answer before you read on.
Two students each draw a random sample of $40$ from the same population and compute the sample mean. One reports $6.1$ , the other reports $6.9$ . Who made the arithmetic error, and how would you decide?
- Compute $\bar{x}$ . Add the sample values, divide by how many there are. One line, and it is the same arithmetic you have done since Year 7.
- Ask what a different sample would have given. This is the move nobody makes unprompted, and it is the entire point of the lesson.
- Define the sample mean for a sample drawn from a population, and compute it.
- Use the correct notation, $\bar{x}$ for the sample mean against $\mu$ for the population mean.
- Recognise that sample means obtained from repeated sampling may be different , even when every sample is the same size.
- Explain why two different sample means are not evidence that anybody made a mistake.
For a sample of $n$ values $x_1, x_2, \ldots, x_n$ drawn from a population, the sample mean is
$$\bar{x} = \frac{x_1 + x_2 + \cdots + x_n}{n} = \frac{1}{n}\sum_{i=1}^{n} x_i$$Read it as: total the sample values, divide by how many you have. The bar over the $x$ is doing important work, and it is not decoration.
| Symbol | Name | Belongs to | Known? |
|---|---|---|---|
| $\bar{x}$ | sample mean | one sample | always, you computed it |
| $\mu$ | population mean | the whole population | generally not |
A worked computation. Five students record their weekly paid work hours: $12, 15, 9, 14, 20$ . Then $n = 5$ and
$$\bar{x} = \frac{12 + 15 + 9 + 14 + 20}{5} = \frac{70}{5} = 14 \text{ hours}$$The sample mean of a sample of n values is x̄ = (x₁ + x₂ + ... + xₙ)/n, that is (1/n)Σxᵢ. It is computed from one sample, so it is always known. The population mean μ is a different quantity and is generally unknown.
Pause, copy the formula in both forms, and write beside it which of $\bar{x}$ and $\mu$ you can always compute and which you generally cannot.
Quick check: A sample of six values totals $92.4$ . What is $\bar{x}$ , to one decimal place?
Here is a population small enough to see all of it, which is a luxury real sampling never gives you. Eight students, and their weekly paid work hours:
| Hours | 1 | 3 | 4 | 6 | 7 | 9 | 10 | 24 |
|---|
The eight values total $64$ , so the population mean is $\mu = 64 \div 8 = 8$ hours. Because this population is tiny we know $\mu$ exactly, which lets us do something you can never do in practice: check the samples against the truth .
Now take four different samples, every one of them of size $n = 4$ :
| Sample | Values | Total | $\bar{x}$ |
|---|---|---|---|
| S1 | $3, 4, 6, 7$ | $20$ | $20 \div 4 = 5$ |
| S2 | $1, 4, 9, 10$ | $24$ | $24 \div 4 = 6$ |
| S3 | $1, 3, 4, 24$ | $32$ | $32 \div 4 = 8$ |
| S4 | $6, 7, 10, 24$ | $47$ | $47 \div 4 = 11.75$ |
Four samples, one population, identical sample size, and four answers: $5$ , $6$ , $8$ and $11.75$ . Nobody made a mistake. Every division above is correct, and every sample is a legitimate sample of size $4$ .
- S3 landed exactly on $\mu = 8$ . That is possible, and it is luck. Nothing about S3's method was better.
- S4 came out at $11.75$ , nearly $50\%$ high , because it happened to catch the extreme value $24$ alongside three moderate values.
- S1 missed the $24$ entirely and came out low at $5$ .
Sample means obtained from repeated sampling may be different, even when every sample is the same size. That variation is not an error: each sample contains different members. One sample mean may land exactly on μ by chance, and from a single sample you cannot tell how close you are.
Pause, copy the four-sample table with its four different means, and beside it the sentence "different samples give different means, and that is not a mistake".
True or false: If two random samples of the same size from the same population give different sample means, at least one of them must contain a calculation error.
Worked examples · 3 in a row, reveal as you go
A random sample of $8$ delivery times, in minutes, is $22, 31, 27, 19, 35, 28, 24, 30$ . Find the sample mean, and state what it estimates. (2 marks)
From the eight-student population above ($1, 3, 4, 6, 7, 9, 10, 24$ hours, $\mu = 8$), sample A is $3, 4, 6, 7$ and sample B is $6, 7, 10, 24$ . Compute both sample means and explain the difference to someone who thinks one of them is wrong. (3 marks)
A random sample of $12$ households has a mean weekly recycling weight of $6.5$ kg. (a) Find the total weight for the sample. (b) A thirteenth household, recycling $19$ kg, is added. Find the new sample mean, and comment. (3 marks)
Complete: A sample of $n$ values has sample mean $\bar{x}$ . The total of the sample values is therefore , and repeated samples of the same size may give sample means.
Misconceptions to fix · the 3 traps that cost marks
True or false: In card 06, sample S3 gave $\bar{x} = 8$ , which equals $\mu$ . This shows S3 was drawn by a better method than the other three samples.
Activities · practice with the ideas
Find the sample mean of $14, 9, 22, 17, 13$ .
A sample of $20$ values has $\bar{x} = 4.85$ . Find the total of the sample values.
From the population $1, 3, 4, 6, 7, 9, 10, 24$ , write down any sample of size $4$ other than the four in card 06, and compute its mean. Is your answer above or below $\mu = 8$ ?
Explain in one sentence why two correct sample means from the same population can differ.
A sample of $9$ has mean $12$ . One value, $30$ , is removed. Find the mean of the remaining $8$ .
Which does NOT belong? True statements about the sample mean $\bar{x}$ :
Earlier you were asked which of two students, reporting $6.1$ and $6.9$ from samples of $40$ , had made the error.
Neither result is necessarily wrong. Two random samples of the same size drawn from the same population contain different members, so they can produce different means. The difference of $0.8$ alone does not establish an error. The question "who was wrong?" assumes $\bar{x}$ is a property of the population. It is a property of the sample. Both numbers may be correct, both estimate the same unknown $\mu$ , and from these two alone you cannot say which is closer to it.
What you can now ask is the useful question: how much do sample means vary, and is $0.8$ a lot? That question has an answer, and the next lesson builds the object that supplies it.
Pick your answer, then rate your confidence. That tells the system what to drill next. Each retry pulls a fresh mix from the bank.
Q1. A random sample of $10$ batteries has lifetimes, in hours, of $42, 51, 47, 39, 55, 44, 48, 50, 46, 38$ . (a) Find the sample mean. (b) State what it estimates, and explain why you cannot assume it equals that quantity. (3 marks)
Q2. Two random samples, each of size $5$ , are drawn from the same population. One has $\bar{x} = 6.2$ and the other $\bar{x} = 7.8$ . Explain why this is not evidence of an error, and state what extra information would be needed to decide which sample mean is closer to $\mu$ . (3 marks)
Q3. A population of eight values is $1, 3, 4, 6, 7, 9, 10, 24$ , with $\mu = 8$ . (a) Give a sample of size $4$ whose mean is below $\mu$ , and one whose mean is above, showing both means. (b) Explain why a sample of size $4$ containing the value $24$ is likely to have a mean above $\mu$ . (c) Give a sample of size $4$ containing $24$ whose mean is not above $\mu$ , and explain what that shows. (4 marks)
Comprehensive answers (click to reveal)
Activity answers:
1. Total $= 14 + 9 + 22 + 17 + 13 = 75$ , $n = 5$ , so $\bar{x} = 75 \div 5 = 15$ .
2. $\sum x_i = n\bar{x} = 20 \times 4.85 = 97$ .
3. Any sample of size $4$ from the eight values, other than the four in card 06. For example $1, 6, 7, 9$ gives $\bar{x} = 23 \div 4 = 5.75$ , which is below $\mu = 8$ . Samples containing $24$ tend to sit above it.
4. Because the two samples contain different members. Each value in a sample of size $n$ contributes $1/n$ of the mean, so swapping even one member changes $\bar{x}$ .
5. Old total $= 9 \times 12 = 108$ . Removing $30$ leaves $78$ across $8$ values, so the new mean is $78 \div 8 = 9.75$ .
Q1 (3 marks): (a) Total $= 460$ , $n = 10$ , so $\bar{x} = 460 \div 10 = 46$ hours [1 for the total and n, 1 for the mean with units]. (b) It estimates $\mu$ , the mean lifetime of the whole population of batteries [1]. You cannot assume $\bar{x} = \mu$ because it was computed from ten batteries rather than the whole population, and a different random sample could give a different value. It could equal $\mu$ by chance, but the sample alone cannot establish that.
Q2 (3 marks): The two samples contain different members, and sample means obtained from repeated sampling may differ even when the samples are the same size, so a gap between them is expected rather than evidence of a mistake [1]. Both are legitimate estimates of the same fixed $\mu$ [1]. To decide which is closer you would need to know $\mu$ itself, and $\mu$ is precisely what is unknown, so nothing computable from the two samples alone can settle it [1].
Q3 (4 marks): (a) Below: $1, 3, 4, 6$ gives $\bar{x} = 14 \div 4 = 3.5$ . Above: $7, 9, 10, 24$ gives $\bar{x} = 50 \div 4 = 12.5$ [1]. (b) In a sample of size $4$ each value carries a weight of $\tfrac{1}{4}$ , and $24$ sits $16$ above $\mu$ , so on its own it lifts the sample mean by $16 \div 4 = 4$ relative to a sample of otherwise typical values. Unless the other three are unusually small, the mean lands above $\mu$ [1 for the weighting argument, 1 for the conclusion]. (c) $1, 3, 4, 24$ gives $\bar{x} = 32 \div 4 = 8$ , exactly $\mu$ . The three smallest values in the population together offset the $24$ completely. This shows the reasoning in (b) is about what is likely across samples, not what is guaranteed for any particular one [1].
A full module quiz covering every lesson in this module, not just this one. Set aside a decent block of time and treat it like a real assessment.
Start the module quiz →Mark lesson as complete
Tick when you've finished the practice and review.