M
hscscience Ext 1 · Y12
0/100daily goal
0
0
0 due
0
L1 · 0 XP
KJ
Your weak spots
Insights load after your first practice round.
Module 10 · Sampling 5 of 5 ~40 min ⚡ +90 XP available

Probabilities for the Sample Mean

Four lessons of build-up, for this. You know the sampling distribution is centred on $\mu$ , that its standard deviation is $\sigma/\sqrt{n}$ , and that for $n \geq 30$ it is approximately normal. Put those together and you can finally answer the question the whole topic exists for: how likely is my sample mean to be close to the truth?

Today's hook, A machine fills bags with a mean of $500$ g and a standard deviation of $80$ g. Sixty-four bags are weighed. Before reading on, write down how likely you think it is that the mean of those $64$ bags lands between $490$ g and $510$ g, and then how likely it is that a single bag does.
0/5QUESTS
01
Recall, your gut answer first

Two numbers, before you read on. Most people put them much closer together than they are.

A machine fills bags with mean $500$ g and standard deviation $80$ g. (a) How likely is the mean of $64$ bags to land between $490$ g and $510$ g? (b) How likely is a single bag to land between $490$ g and $510$ g?

auto-saved
02
The two moves, and they never change
  1. Find $\text{sd}(\bar{X}) = \dfrac{\sigma}{\sqrt{n}}$ , and check $n \geq 30$ . Without that check you have no right to the normal curve at all.
  2. Standardise each bound and read the area. Convert the given values of $\bar{x}$ into $z$ , then use the empirical rule to turn the $z$ range into a probability.
One symbol changes, and it is the whole lesson. Standardising an individual value divides by $\sigma$ . Standardising a sample mean divides by $\sigma/\sqrt{n}$ . Everything else about the process is identical, which is exactly why the wrong one is so easy to reach for.
03
What you'll master
  • Standardise a sample mean with $z = \dfrac{\bar{x} - \mu}{\sigma/\sqrt{n}}$ .
  • Use the empirical rule to estimate the probability that $\bar{X}$ lies within given bounds .
  • Handle the three shapes: between two bounds, above a bound, below a bound.
  • Explain why the same bounds are far more likely for a mean than for one observation.
04
Key terms
Standardising $\bar{x}$Converting a sample mean into a $z$ value by subtracting $\mu$ and dividing by $\sigma/\sqrt{n}$ . Like this: $\bar{x} = 510$ with $\mu = 500$ and $\sigma/\sqrt{n} = 10$ gives $z = 1$ .
Empirical ruleThe areas $68\%$ , $95\%$ and $99.7\%$ within $1$ , $2$ and $3$ standard deviations of the mean. Like this: $-1 < z < 1$ encloses $68\%$ .
TailThe area beyond a bound, on one side. Like this: beyond $z = 2$ is $2.5\%$ , because $5\%$ sits outside $\pm 2$ and symmetry splits it.
Within given boundsThe probability that $\bar{X}$ falls between two stated values. Like this: $P(490 < \bar{X} < 510)$ .
05
Standardising a sample mean
core concept

By the central limit theorem, for $n \geq 30$ the sample mean is approximately $N\!\left(\mu, \dfrac{\sigma^2}{n}\right)$ . To find a probability you convert $\bar{x}$ to a $z$ value in the usual way, using that distribution's standard deviation:

$$z = \frac{\bar{x} - \mu}{\dfrac{\sigma}{\sqrt{n}}}$$

Then read the area from the empirical rule:

$z$ rangearea insidearea outsideeach tail
$-1$ to $1$$68\%$$32\%$$16\%$
$-2$ to $2$$95\%$$5\%$$2.5\%$
$-3$ to $3$$99.7\%$$0.3\%$$0.15\%$

The hook, worked. Bags with $\mu = 500$ , $\sigma = 80$ , and $n = 64$ :

  • $\text{sd}(\bar{X}) = \dfrac{80}{\sqrt{64}} = \dfrac{80}{8} = 10$ g, and $n = 64 \geq 30$ , so the theorem applies.
  • $490$ standardises to $z = \dfrac{490 - 500}{10} = -1$ , and $510$ to $z = +1$ .
  • So $P(490 < \bar{X} < 510) = P(-1 < z < 1) = \mathbf{68\%}$ .
Why $\sigma/\sqrt{n}$ and not $\sigma$ . You are asking about the behaviour of an average of $64$ bags , and averages vary far less than individual bags do. Dividing by $\sigma$ would answer a question about one bag, and would give an answer roughly eight times too pessimistic here, because $\sqrt{64} = 8$ .

To find a probability for a sample mean, first check n ≥ 30 and compute sd(X̄) = σ/√n, then standardise each bound with z = (x̄ − μ) ÷ (σ/√n) and read the area from the empirical rule: 68% within ±1, 95% within ±2, 99.7% within ±3.

Pause, copy the $z$ formula for a sample mean with the $\sigma/\sqrt{n}$ clearly in the denominator, and copy the empirical-rule table including the tail column.

Quick check: With $\mu = 500$ , $\sigma = 80$ and $n = 64$ , what is $z$ for $\bar{x} = 520$ ?

06
The three shapes, and one comparison worth remembering
core concept

Every question of this kind is one of three shapes, and each is the same two moves followed by a different reading. Staying with the bags, where $\text{sd}(\bar{X}) = 10$ :

AskedStandardisedRead asAnswer
$P(480 < \bar{X} < 520)$$P(-2 < z < 2)$the area inside $\pm 2$$95\%$
$P(\bar{X} > 520)$$P(z > 2)$half of the $5\%$ outside$2.5\%$
$P(\bar{X} < 480)$$P(z < -2)$the other half, by symmetry$2.5\%$

Two facts do all the work in the second and third rows: the curve is symmetric , and the total area is $1$ . A one-sided question is always the inside area subtracted from $1$ , then halved if symmetry applies.

Now the comparison the hook was really about. Ask the same question about a single bag rather than the mean of $64$ :

Between $490$ g and $510$ gdivide by$z$ rangeprobability
the mean of $64$ bags$\sigma/\sqrt{n} = 10$$-1$ to $1$about $68\%$
one single bag$\sigma = 80$$-0.125$ to $0.125$about $10\%$
Same bounds, same population, and nearly seven times the probability. For one bag, $490$ to $510$ is a sliver barely an eighth of a standard deviation wide. For the mean of $64$ , the same interval is a full standard deviation either side. Averaging did not change the bags. It changed how much the thing you are measuring moves around. That is the payoff of the whole arc, and it is why a sample of $64$ can say something confident about a population nobody has measured.

Between two bounds, read the inside area. Above or below a bound, subtract the inside area from 1 and halve it if the curve's symmetry applies. The same bounds are far more likely for a sample mean than for one observation, because the mean's spread is σ/√n rather than σ.

Pause, copy the three shapes with their readings, and copy the two-row comparison showing $68\%$ against $10\%$ for the same interval.

True or false: Because the population has $\sigma = 80$ , the probability that the mean of $64$ bags lies within $10$ g of $500$ g is the same as the probability that one bag does.

PROBLEM 1 · BETWEEN TWO BOUNDS

Test scores have $\mu = 72$ and $\sigma = 15$ . A random sample of $25$ scores is taken. Estimate $P(66 < \bar{X} < 78)$ . (3 marks)

1
$\text{sd}(\bar{X}) = \dfrac{15}{\sqrt{25}} = \dfrac{15}{5} = 3$
Always compute this first. It is the number every later step divides by, and using 15 here is the standard error of this whole topic.
PROBLEM 2 · ONE TAIL

Bags have $\mu = 500$ g and $\sigma = 80$ g. For a sample of $64$ bags, estimate the probability that the sample mean exceeds $520$ g. (2 marks)

1
$\text{sd}(\bar{X}) = \dfrac{80}{\sqrt{64}} = 10$ , and $n = 64 \geq 30$
Condition met, so the normal approximation is licensed without any assumption about the population's shape.
PROBLEM 3 · TURN IT INTO A COUNT

For the same bags, $200$ separate samples of $64$ are taken and each sample mean recorded. (a) About how many of those $200$ means would you expect to fall between $470$ g and $530$ g? (b) About how many below $480$ g? (3 marks)

1
$470$ and $530$ give $z = \dfrac{\pm 30}{10} = \pm 3$
Same sd(X̄) = 10 as before; only the bounds changed.

Complete: To standardise a sample mean, divide by , and the area within $\pm 2$ of the mean is %.

Trap 01
Standardising with $\sigma$ instead of $\sigma/\sqrt{n}$
The most expensive error in the topic, and the hardest to spot, because the answer still looks like a probability. With the bags it turns $z = 1$ into $z = 0.125$ and $68\%$ into about $10\%$ . If the question says "sample mean", the denominator has a $\sqrt{n}$ in it.
Trap 02
Quoting the whole outside area for one tail
Beyond $\pm 2$ lies $5\%$ , but $P(\bar{X} > \text{upper bound})$ is only half of it, $2.5\%$ . Sketch the curve and shade what was asked for: one tail or two is visible in the picture and invisible in the algebra.
Trap 03
Not checking $n \geq 30$ , or checking it silently
The normal approximation has a condition. When $n \geq 30$ , say so, because it is the justification. When $n < 30$ , as in worked example 1, you may still proceed but you must state the extra assumption that the population is roughly normal. Marks are awarded for the condition, not only the arithmetic.

True or false: If $5\%$ of the area lies outside $z = \pm 2$ , then $P(\bar{X} > \mu + 2\,\text{sd}) = 5\%$ .

1

A population has $\mu = 200$ and $\sigma = 24$ . For $n = 36$ , find $\text{sd}(\bar{X})$ , then estimate $P(196 < \bar{X} < 204)$ .

2

For the same population and sample size, estimate $P(\bar{X} > 208)$ .

3

Explain in one sentence why $P(196 < \bar{X} < 204)$ is much larger than the probability that a single observation falls between $196$ and $204$ .

4

$400$ samples of size $36$ are taken from that population. About how many sample means would you expect below $192$ ?

5

A student standardises a sample mean by dividing by $\sigma$ . Describe the effect on their answer, and how you would spot it in their working.

Which does NOT belong? Steps in estimating $P(a < \bar{X} < b)$ :

11
Revisit your thinking

Earlier you estimated two probabilities for the same interval, $490$ g to $510$ g.

For the mean of $64$ bags it is about $68\%$ , because $\text{sd}(\bar{X}) = 80/\sqrt{64} = 10$ and the interval is exactly one standard deviation either side. For a single bag it is about $10\%$ , because that same interval is only $\pm 0.125$ of the population's own standard deviation of $80$ . Nearly seven times the probability, from the identical population and the identical bounds.

That gap is the whole of Module 10's second half in one number. Lesson 21 said a sample can stand in for a population it cannot measure. Lesson 22 showed sample means vary. Lesson 23 gave that variation a distribution, Lesson 24 gave it a shape and a width, and this lesson turned it into a probability you can act on. ME1-12-06 is complete.

auto-saved
01
Multiple choice
+5 XP per correct · +25 XP all-correct

Pick your answer, then rate your confidence. That tells the system what to drill next. Each retry pulls a fresh mix from the bank.

02
Short answer
ApplyBand 33 marks

Q1. A population has $\mu = 150$ and $\sigma = 20$ . A random sample of $100$ is taken. (a) Find $\text{sd}(\bar{X})$ . (b) Estimate $P(146 < \bar{X} < 154)$ , justifying your use of the normal approximation. (3 marks)

auto-saved
ApplyBand 43 marks

Q2. For the same population and sample size, estimate $P(\bar{X} > 156)$ . Show how you obtain a one-tailed probability from the empirical rule. (3 marks)

auto-saved
AnalyseBand 54 marks

Q3. Still with $\mu = 150$ , $\sigma = 20$ , $n = 100$ . (a) Five hundred separate samples are taken and each sample mean recorded. About how many would you expect below $144$ ? (b) A student standardises using $\sigma = 20$ instead of $\text{sd}(\bar{X})$ . State the factor by which their $z$ values are wrong, and say whether their probabilities come out too large or too small. (c) Explain what would change in your answer to (a) if the sample size were $9$ rather than $100$ . (4 marks)

auto-saved
Comprehensive answers (click to reveal)

Activity answers:

1. $\text{sd}(\bar{X}) = \dfrac{24}{\sqrt{36}} = 4$ . Then $196$ and $204$ give $z = -1$ and $z = 1$ , so $P \approx 68\%$ .

2. $z = \dfrac{208 - 200}{4} = 2$ . The area inside $\pm 2$ is $95\%$ , so $5\%$ lies outside, and by symmetry one tail is $\mathbf{2.5\%}$ .

3. Because the sample mean varies far less than a single observation: its standard deviation is $\sigma/\sqrt{n} = 4$ rather than $\sigma = 24$ . The same interval is therefore $\pm 1$ standard deviation for the mean but only about $\pm 0.17$ for one observation.

4. $z = \dfrac{192 - 200}{4} = -2$ , so $P(\bar{X} < 192) = 2.5\%$ , and $0.025 \times 400 = \mathbf{10}$ sample means.

5. Dividing by $\sigma$ instead of $\sigma/\sqrt{n}$ makes every $z$ value $\sqrt{n} = 6$ times too small, so the probability of being close to $\mu$ comes out far too low. You would spot it in the working as a denominator with no $\sqrt{n}$ in it, and in the answer as an implausibly small probability for a large sample.

Q1 (3 marks): (a) $\text{sd}(\bar{X}) = \dfrac{20}{\sqrt{100}} = \dfrac{20}{10} = 2$ [1]. (b) $146$ gives $z = \dfrac{146 - 150}{2} = -2$ and $154$ gives $z = 2$ [1], so $P(-2 < z < 2) \approx 95\%$ . The normal approximation is justified by the central limit theorem because $n = 100 \geq 30$ , and no assumption about the population's shape is needed [1].

Q2 (3 marks): $z = \dfrac{156 - 150}{2} = 3$ [1]. The empirical rule gives $99.7\%$ inside $\pm 3$ , so $0.3\%$ lies outside [1]. The curve is symmetric, so one tail is half of that: $P(\bar{X} > 156) \approx \mathbf{0.15\%}$ [1].

Q3 (4 marks): (a) $144$ gives $z = -3$ , so $P(\bar{X} < 144) \approx 0.15\%$ , and $0.0015 \times 500 = 0.75$ , so about $1$ of the $500$ sample means [1]. (b) Dividing by $20$ instead of $2$ makes every $z$ value $10$ times too small, that is $\sqrt{n}$ times [1]. A smaller $z$ puts the bound nearer the centre, so the tail probabilities come out far too large and the "within bounds" probabilities far too small [1]. (c) With $n = 9$ , $\text{sd}(\bar{X})$ would be $\dfrac{20}{3} \approx 6.67$ rather than $2$ , so $144$ would be less than one standard deviation below the mean and the probability would be much larger. More importantly $n = 9 < 30$ , so the central limit theorem no longer licenses the normal approximation, and the estimate would only be defensible if the population were already known to be roughly normal [1].

01
Take the full module quiz
quiz

A full module quiz covering every lesson in this module, not just this one. Set aside a decent block of time and treat it like a real assessment.

Start the module quiz →

Mark lesson as complete

Tick when you've finished the practice and review.