Populations and Samples
Everything in this module so far assumed you already knew $p$ . Real statistics almost never hands you $p$ , or the mean, or the shape of the distribution. You get a handful of measurements and have to say something defensible about a group you will never measure in full. This lesson sets up the two objects that whole argument rests on, the population and the sample , and the conditions under which the second tells you anything about the first.
Write your answer before you read on. Being wrong here is useful; being vague is not.
A news site runs an online poll. $8{,}000$ readers click, $62\%$ choose option A, and the headline reads "Australians prefer A". What would have to be true about those $8{,}000$ people for that headline to be defensible? And does $8{,}000$ being a large number help?
Sampling questions look wordy and open-ended. They are not. Almost all of them are answered by making two moves in order.
- Name the population exactly. Who or what, and all of them. The population is fixed by the question being asked, never by who happened to be available.
- Interrogate the sample on two separate counts. Is it random ? Is $n$ large enough , typically $n \geq 30$ ? These are two different questions with two different jobs, and a sample can pass one and fail the other.
- Define a statistical population as the entire group of people or objects about which information is sought.
- Define a sample as a selection of people or objects drawn from a population.
- Recognise that a statistic taken from a sample approximates the population well when the sample is random and $n$ is large enough , typically $n \geq 30$ .
- Recognise that, in general, both the distribution of a population and the distribution of a statistic such as its mean are unknown .
- Recognise that the distribution of a random sample is likely to resemble the distribution of the population once $n$ is large enough.
A statistical population is the entire group of people or objects about which information is sought. Two phrases in that definition do real work, and both are examinable.
- "Entire group." Not most of it, not the reachable part of it. The population is the whole thing, which is precisely why you usually cannot measure it.
- "About which information is sought." The population is defined by the question . Change the question and you change the population, even if the people in the room stay the same.
A sample is a selection of people or objects drawn from that population. The sample is the part you actually measure, and it exists only because the population is out of reach.
Once the two are named, every number in the problem belongs to exactly one of them:
| Population | Sample | |
|---|---|---|
| What it is | the entire group | a selection drawn from it |
| Its numbers are called | parameters | statistics |
| Mean | $\mu$ | $\bar{x}$ |
| Variance | $\sigma^2$ | $s^2$ |
| Can you compute it? | generally no | always yes |
| Does it change? | no, it is one fixed number | yes, a new sample gives a new value |
A statistical population is the entire group of people or objects about which information is sought; a sample is a selection drawn from it. Population numbers are parameters (μ, σ²), fixed and usually unknown. Sample numbers are statistics (x̄, s²), always computable and different for every sample.
Pause, copy the population-versus-sample table, and beside it write the sentence "the population is decided by the question, not by who was available".
Quick check: A council wants to know the average weekly water use of the $12{,}400$ households in its area, and meters $200$ of them. What is the population?
Start with the admission that makes sampling necessary at all.
Given that, when does a sample statistic actually approximate the population parameter? Two conditions, and they are independent of each other:
- The sample is random. Selection must not depend on the thing you are measuring.
- The sample size is large enough , typically $n \geq 30$ .
Meet both and two useful things follow: the statistic is a good approximation of the parameter, and the shape of the sample's distribution is likely to resemble the shape of the population's distribution. A random sample of $60$ incomes from a right-skewed population will itself usually look right-skewed.
The two conditions do different jobs, and this is the examinable point:
| Condition | What it protects against | What goes wrong without it |
|---|---|---|
| Random | bias, a systematic error | the sample is centred in the wrong place |
| $n \geq 30$ | variability, ordinary noise | the sample bounces around too much to be useful |
A sample statistic approximates the population parameter well when the sample is random AND n is large enough, typically n ≥ 30. Random controls bias; large n controls variability. They are independent, and a larger n never fixes a biased sample. In general both the population distribution and the distribution of a statistic are unknown.
Pause, copy the two conditions with the job each one does, and copy the sentence "a bigger sample does not fix bias, it only makes a wrong answer more precise".
True or false: A survey of $5{,}000$ people who volunteered to take part will estimate a population mean more accurately than a random sample of $50$ , because $5{,}000$ is much larger than $50$ .
Worked examples · 3 in a row, reveal as you go
A school has $1{,}240$ students. To estimate the mean daily screen time of its students, $40$ names are drawn at random from the full roll and surveyed. Their mean daily screen time is $3.4$ hours. Identify the population, the sample, the sample size, the parameter of interest and the statistic.
Two attempts are made to estimate the mean daily screen time of the same $1{,}240$ students. Sample P: $40$ students drawn at random from the roll. Sample Q: $800$ students who responded to a notice posted in the school's gaming club channel. Which sample gives the better approximation, and why? (3 marks)
Weekly household income in a large city is strongly right-skewed: most households cluster at lower incomes with a long tail of high earners. A researcher draws a random sample of $60$ households, and a second researcher draws a random sample of $6$ . What can be said about the shape of each sample's distribution, and about what remains unknown?
Complete: A sample statistic is a good approximation of a population parameter when the sample is and the sample size is at least .
Misconceptions to fix · the 3 traps that cost marks
True or false: Because a random sample of size $n = 45$ was taken, the population mean $\mu$ is now known.
Activities · practice with the ideas
A quality inspector wants the mean lifetime of the $50{,}000$ batteries in a production run, and tests $80$ of them. Name the population, the sample, $n$ , the parameter and the statistic.
A researcher surveys shoppers outside a supermarket on a Tuesday morning to estimate the mean weekly grocery spend of all households in the suburb. Give one reason this sample may not be random, and state the likely direction of the bias.
Explain, in one sentence each, the different job done by randomness and by a large $n$ .
A population of reaction times is right-skewed. A random sample of $n = 100$ is drawn. What can you say about the shape of the sample, and what can you still not say about the population mean?
A student writes that because 2000 people were sampled, the estimate must be accurate. Rewrite this as a statistically defensible sentence, adding whatever condition is missing.
Which does NOT belong? Things that are true of a statistical population:
Earlier you judged a headline built on $8{,}000$ online clicks.
The $8{,}000$ are self-selected: they read that site, saw that poll, and chose to click. Selection depends on exactly the opinions being measured, so the sample is biased, and no sample size repairs that. The poll does estimate something precisely, the views of people who choose to click polls on that site, but that group is not "Australians". The honest headline names its population. And note what $8{,}000$ did buy: very little variability. The result is repeatable and stable, and still wrong, which is what makes large biased samples so persuasive.
Pick your answer, then rate your confidence. That tells the system what to drill next. Each retry pulls a fresh mix from the bank.
Q1. A gym has $3{,}400$ members and wants to estimate the mean number of visits per member per month. Fifty members are selected at random from the membership list, and their mean is $6.2$ visits. State the population and the statistic, and explain whether both conditions for a good approximation are met. (3 marks)
Q2. A company emails a satisfaction survey to all $20{,}000$ of its customers, and $1{,}500$ reply. Explain why the $1{,}500$ may not be a random sample, state the likely direction of any bias, and explain why raising the number of replies to $5{,}000$ would not remove the problem. (3 marks)
Q3. Explain why, in general, both the distribution of a population and the distribution of a statistic such as the sample mean are unknown. State the two conditions under which a sample statistic is a good approximation of a population parameter, and explain what each condition protects against. (4 marks)
Comprehensive answers (click to reveal)
Activity answers:
1. Population: all $50{,}000$ batteries in the production run. Sample: the $80$ tested. $n = 80$ . Parameter: $\mu$ , the mean lifetime of all $50{,}000$ , unknown. Statistic: the mean lifetime of the $80$ tested, computable.
2. Shoppers outside one supermarket on a Tuesday morning are not a random selection of all households in the suburb: people who shop on a weekday morning are more likely to be retired or not in full-time work, and households that shop elsewhere or online cannot be selected at all. Selection is not independent of grocery habits, so the estimate is likely to sit below the true mean weekly spend if those households buy in smaller, more frequent trips. Naming a direction and a reason is what earns the second mark.
3. Randomness protects against bias , a systematic error that puts the sample's centre in the wrong place. A large $n$ protects against variability , the ordinary noise that makes a small sample bounce around. They are independent, and only the first is about which members were chosen.
4. The sample is random with $n = 100 \geq 30$ , so it is likely to resemble the population and should itself appear right-skewed. About $\mu$ you can still say only that the sample mean estimates it: $\mu$ remains unknown, and a different sample of $100$ would give a different estimate.
5. "We took a random sample of $2{,}000$ people, and since the sample was random and $n \geq 30$ , the sample mean is a good approximation of the population mean." The missing condition is randomness; size alone establishes nothing.
Q1 (3 marks): Population: all $3{,}400$ gym members [1]. Statistic: $\bar{x} = 6.2$ visits per month, the mean of the $50$ sampled members [1]. Both conditions are met: members were selected at random from the full membership list, so selection does not depend on how often they visit, and $n = 50 \geq 30$ [1]. So $6.2$ is a good approximation of $\mu$ , though it does not equal it.
Q2 (3 marks): The $1{,}500$ are self-selected: they chose to reply, and customers with strong opinions, particularly the very satisfied or the very dissatisfied, are more likely to do so [1]. The likely direction is an overstatement of extreme satisfaction ratings relative to the full customer base, since indifferent customers are least likely to reply [1]. Raising replies to $5{,}000$ changes $n$ but not the selection mechanism, so the bias remains; the estimate simply becomes a more precise estimate of the wrong quantity [1].
Q3 (4 marks): The population distribution is unknown because the entire group is not measured, which is the reason for sampling in the first place; $\mu$ and $\sigma$ are therefore unknown [1]. The distribution of the sample mean is also unknown, because $\bar{x}$ varies from sample to sample and its behaviour depends on the very population distribution that is unknown [1]. A sample statistic is a good approximation when the sample is random and when $n$ is large enough, typically $n \geq 30$ [1]. Randomness protects against bias, a systematic error in which members are selected in a way that depends on what is being measured; a large $n$ protects against variability, so that ordinary sampling noise is small. A larger $n$ does not correct bias [1].
A full module quiz covering every lesson in this module, not just this one. Set aside a decent block of time and treat it like a real assessment.
Start the module quiz →Mark lesson as complete
Tick when you've finished the practice and review.