Sampling Methods and Sample Size
Two different questions get confused constantly. How the sample was chosen decides whether it represents the population at all. How big it is decides only how precise the answer is. A bad method is not repaired by a bigger sample.
You want to know the mean daily screen time of students at your school. You stand outside the library at lunchtime and ask the first fifty students who come out. Write down two groups of students who are unlikely to appear in your sample, and say whether you would expect their screen time to be higher or lower than average.
A sample is only useful if it is representative: every part of the population must have a genuine chance of being in it. The way the sample is chosen decides that. The size of the sample decides something different and narrower: how much the answer would wobble if you did the whole thing again.
Method decides whether the answer is about the right people. Size decides how precise it is.
The two are not interchangeable, and the direction matters. A well-chosen small sample gives an imprecise answer to the right question. A badly chosen large sample gives a precise answer to the wrong question, which is far more dangerous, because precision is persuasive.
Know
- The common sampling methods: random, systematic, stratified, and convenience
- That bias comes from the method and imprecision comes from the size
- That a self-selected sample is not representative however large it is
Understand
- Why increasing the size of a biased sample makes the wrong answer more precise, not more correct
- Why non-response can bias a sample that was selected properly
Can Do
- Identify the sampling method used in a report and name who it excludes
- Choose an appropriate method for a given population and aim
- Judge whether a stated conclusion is supported by the sample described
Each has a cost and a failure mode.
| Method | How | Fails when |
|---|---|---|
| Random | every member has an equal chance, chosen by a list and a random number | you have no complete list of the population |
| Systematic | take every $k$th member of an ordered list | the list has a repeating pattern that lines up with $k$ |
| Stratified | split into subgroups, sample each in proportion | the subgroups are wrongly defined or their sizes are unknown |
| Convenience | whoever is easy to reach | almost always; "easy to reach" is itself a characteristic |
Stratified sampling deserves attention because it is the one that fixes a specific known problem. If a school is 55% girls and 45% boys and you sample 40 students, a stratified sample takes 22 and 18 rather than trusting chance to balance them. It removes one source of variation you already know about.
Bias is a systematic error: the sample is not a small copy of the population but a copy of some particular part of it. There is one diagnostic question, and it is worth asking before anything else.
Who could not have been selected?
Survey students outside the library at lunch and you have excluded everyone who eats elsewhere, everyone at sport, and everyone absent. If those groups differ from the library group in the thing you are measuring, and they usually do, your answer is wrong in a fixed direction.
The 1936 example is the definitive one. Ballots were mailed to people on lists of car owners and telephone subscribers, so in a depression the sample was systematically wealthier than the electorate, and wealth was related to voting intention. Ten million ballots could not repair that, because every additional ballot came from the same skewed list.
Size buys precision, and only precision. If you drew a fresh random sample tomorrow you would get a slightly different mean; the size controls how different.
The relationship has a shape worth knowing: precision improves with the square root of the sample size, not with the size itself. To halve the wobble you must quadruple the sample.
| Sample size | Relative wobble |
|---|---|
| $100$ | $1$ |
| $400$ | $\tfrac{1}{2}$ |
| $1600$ | $\tfrac{1}{4}$ |
Two consequences follow. Going from $100$ to $400$ is worth a great deal; going from $1600$ to $1700$ is worth almost nothing. And a national poll of about $1000$ people can be genuinely informative, which surprises people, because $1000$ well-chosen responses already give a small wobble regardless of whether the population is a school or a country.
A sample can be chosen perfectly and still end up biased, if the people who decline differ from the people who reply.
Suppose you select 100 students at random and 30 reply. You do not have a random sample of 100; you have a self-selected sample of 30. The question becomes: who bothers to reply? People with strong opinions, people with time, people who like the researcher.
If you are surveying satisfaction with the canteen, the people who reply are disproportionately the people who feel strongly, and strong feelings about a canteen skew negative. The result is not a small random error but a predictable shift.
Watch Me Solve It · 3 examples
-
1Name the methodThis is a convenience sample. The selection criterion is "present at a particular place and time", which is a characteristic of the students, not a random choice.
-
2Ask who could not have been selectedStudents at sport, students eating elsewhere, students absent, and students who never use the library. None of these had any chance of selection.
-
3Predict the direction of the biasLibrary users at lunch are plausibly more study-oriented, so their screen time may be lower than average. The sample mean would therefore be an under-estimate, not merely an imprecise estimate.
-
4State what would fix it, and what would notA random selection from the full school roll would fix it. Asking 500 students at the library instead of 50 would not: every extra response comes from the same excluded-heavy group.
-
1Find the total and the proportions$300 + 240 + 180 = 720$The strata are the year groups, and each must be represented in proportion to its size.
-
2Compute each stratum's share$\frac{300}{720} \times 60 = 25, \quad \frac{240}{720} \times 60 = 20, \quad \frac{180}{720} \times 60 = 15$The three parts sum to 60, as they must.
-
3Sample randomly within each stratum$25 \text{ from Year 9}, \quad 20 \text{ from Year 10}, \quad 15 \text{ from Year 11}$Randomness within each group is still required; stratifying replaces chance only for the year-group split.
-
4State the advantageA simple random sample of 60 could by chance contain very few Year 11 students. Stratifying guarantees the year proportions match the school, removing a source of variation that is already known about.
-
1Identify the population claimed and the population sampledThe claim is about all Australians. The sample is visitors to one website who chose to click a poll: a different population entirely.
-
2Name both problemsConvenience sampling, since only that site's visitors could be selected; and self-selection, since only those motivated to click were counted.
-
3Assess what the size contributes12,000 makes the figure precise for the group that responded, and contributes nothing to whether that group resembles Australians. The precision makes the wrong answer more persuasive.
-
4State what could honestly be claimedAt most: "78% of the 12,000 visitors to this site who chose to respond expressed support." That is a defensible sentence, and it is much weaker than the one printed.
Brain Trainer · 4 problems
Four quick problems. Work each one, then reveal the answer.
-
1 A survey of 5000 people is taken by stopping shoppers at one mall. Is it representative?
Everyone who does not visit that mall had no chance of selection.No, it is a convenience sample -
2 To halve the wobble in an estimate, what must happen to the sample size?
Precision improves with the square root of the size.It must quadruple -
3 100 students are randomly selected and 22 reply. What should be reported?
A reader cannot judge non-response without both numbers.The response rate, 22 of 100 -
4 A school is 60% girls. A stratified sample of 50 takes how many girls?
$0.6 \times 50$.$30$
Multiple Choice · 5 questions
A sampling method that systematically excludes part of the population is best repaired by:
Increasing a sample from $100$ to $400$ changes the wobble in the estimate by a factor of about:
An online poll on a news site attracts 50,000 responses. The main problem is that:
A school of 800 students is 45% in the junior years. A stratified sample of 80 should contain how many junior students?
300 people are randomly selected and 60 reply. The 60 replies are:
Short Answer · 3 questions
(a) Surveying students who attend a voluntary after-school study session, to estimate homework hours.
(b) Selecting every 10th name from an alphabetical school roll.
(c) Posting a survey link in the school newsletter and using whoever responds.
(a) Explain what surveying 300 students would improve.
(b) Explain what it would NOT improve, if the 300 were also friends and friends-of-friends.
(c) Describe a method that would improve what the larger sample cannot, and explain why it works.
(a) State the population the claim is about and the population actually sampled.
(b) Name two distinct problems with the sampling.
(c) Rewrite the headline so that it is supported by the data described.
(a) The ballots were mailed to names drawn from car registrations and telephone directories. Explain precisely how this produced the error.
(b) A response rate of 24% was also involved. Explain how that compounds the problem.
(c) The rival's method has since been criticised too. Suggest what could go wrong with matching quotas to known population proportions, and how modern polls try to address it.
Method
Decides representativeness. Ask who was excluded
Size
Decides precision only, by the square root
Self-selection
Not random, however many respond
Report
The response rate, not just the count
Your Badges
0 of 6Mark lesson as complete
Tick when you've finished Learn, Practice and the Stretch. Earns +85 XP and +25 coins.