Year 12 Biology Module 8 · IQ3 ⏱ ~45 min Practice bank · 3 Short Answer Lesson 13 of 21 Data analysis focus

Evaluating Epidemiological Study Methods

A study's conclusion is only as strong as its method. Learn how design, sampling, bias, confounding and follow-up affect what epidemiological evidence can prove.

Today's hook: Two studies report the same disease pattern, but one is a small survey and the other follows thousands of people over ten years. Should you trust them equally?
0/5TASKS
1
You’re here

Get oriented

Warm up first

Three quick questions from earlier lessons. Pulling old material back to mind before you learn something new makes the new material stick better, so this is not busywork.

Worksheets

Practise this lesson

Four printable worksheets that build from the foundations up to exam-style questions, start at whatever level suits you.

Lesson map

Method -> evidence -> judgement

Use a study checklist to decide whether a conclusion is well supported.

  1. Identify the design.Different study designs answer different questions.
  2. Test the method.Look for sample size, control groups, bias and confounding.
  3. Judge the claim.State what the method supports and what it cannot prove.

Know what matters

Must Know
  • Study method affects the strength of epidemiological evidence.
  • Bias is a systematic error that can distort results.
  • Confounding occurs when another factor partly explains the pattern.
  • Evaluation needs a judgement, not just a description.
Should Know
  • Cohort, case-control, cross-sectional and RCT designs have different strengths.
  • Blinding, controls and follow-up can improve study quality.
  • Statistical significance is not the same as practical importance.
Going Deeper
  • Absolute risk reduction, relative risk reduction and NNT.
  • Systematic reviews and replication.
  • Survival curves and endpoint selection.
0
Predict first: what weakens the claim?
connect

A survey finds people who eat breakfast report lower heart disease rates. Which issue most directly limits a causal claim?

5

Key vocabulary, translated

1
Key vocabulary, translated
vocab
MethodEverything about how the study was actually run: who was studied, what was measured, how long they were followed, and what they were compared against.Like this: 1,000 smokers and 1,000 non-smokers, lung cancer diagnoses counted from hospital records, tracked for 10 years.
BiasA flaw built into the study that tips every result the same way. Because it is baked into the design, collecting more data does not cancel it out.Like this: asking people to recall their diet from 20 years ago. People who are now sick search their memory harder for a cause, so their reported junk food intake comes out too high.
ConfounderA second difference between the groups that could be the real cause of the pattern, so the exposure you are studying gets blamed for damage something else did.Like this: coffee drinkers get more lung cancer, but coffee drinkers also smoke more. Smoking is the confounder doing the damage.
Control groupThe group that shows what would have happened anyway. It is matched to the test group in every way except the one thing being tested.Like this: in a vaccine trial the control group gets a saline injection, so the only difference between the two groups is the vaccine itself.
Sample sizeHow many people or records the study used. A bigger sample squeezes out random luck, but it cannot repair a study that was designed badly.Like this: a pattern in 12 patients could easily be chance. The same pattern in 40,000 probably is not, unless bias created it.

True or false: a large sample automatically removes all bias.

2
What each design is good for
apply

Cohort

Follows people over time. Strong for showing exposure before disease.

Case-control

Starts with people who have or do not have the disease. Useful for rare diseases.

RCT

Randomly allocates a treatment or prevention strategy. Strong for testing interventions.

Cross-sectional studies take a snapshot at one time. They are useful for prevalence, but weaker for proving which factor came first.

Remember!

Evaluate a method by naming the design, stating a strength, identifying a limitation, and making a justified judgement.

Pause and copy this four-part frame before sorting the steps.

Build an evaluation+7 XP

Put the study-evaluation steps in order.

  • Identify one limitation such as bias, confounding or short follow-up.
  • Name the study design.
  • Judge whether the conclusion is strong, limited or not supported.
  • State one strength of the method.
3
The evidence hierarchy: ranking studies by trustworthiness
classify

Not all evidence is equal. A single patient case report can flag a new phenomenon, but it cannot establish a general truth. A well-conducted systematic review that pools dozens of randomised trials gives the most reliable basis for a clinical decision. Epidemiologists rank study designs in a hierarchy, and a design's position depends on how well it controls confounding and bias.

At the top sits the systematic review with meta-analysis of randomised controlled trials, because pooling many trials averages out chance variation and single-study quirks; a Cochrane review of statin trials is the classic example. One level down, a single well-designed RCT still establishes causation through random allocation, though it may not generalise. The UKPDS trial of metformin for type 2 diabetes shaped diabetes care for decades.

Cohort studies follow exposed and unexposed groups forward in time, so they establish that the exposure came before the disease; the Nurses' Health Study tracking diet and cancer is a landmark. Case-control studies work backwards from people with and without a disease, efficient for rare conditions such as HPV and cervical cancer, but vulnerable to recall bias. Cross-sectional surveys give cheap snapshots, and case reports sit at the base.

The hierarchy explains why media claims so often mislead. A headline built on one cross-sectional survey sits near the bottom of the pyramid, yet it is reported with the confidence of a meta-analysis. In an exam, naming where a study sits in the hierarchy, and justifying that position, is an instant evaluation mark.

Book notes
  • Hierarchy, strongest to weakest: systematic review with meta-analysis, RCT, cohort, case-control, cross-sectional, case report.
  • Position depends on how well a design controls confounding and bias.
  • Case-control suits rare diseases but suffers recall bias; cross-sectional gives prevalence snapshots only.
  • Exam move: name the study's level in the hierarchy and justify it.

Which source provides the strongest evidence for a causal claim about a treatment?

7

From data to a fair claim

4
Three numbers that test any treatment claim
explain

In 1998 the NSABP P-1 trial randomised 13,388 women at high risk of breast cancer to tamoxifen or placebo. The tamoxifen group developed 89 cases of invasive breast cancer; the placebo group developed 175. Press releases announced a 49% reduction, yet the same data gives an absolute risk reduction of 0.64 percentage points and a number needed to treat of 156. All three figures are arithmetically correct.

Worked example: a statin trial

Take a statin trial that follows 10,000 patients for five years: 100 heart attacks among 5,000 on the statin (risk 0.02) and 150 among 5,000 on placebo (risk 0.03). Relative risk is 0.02 divided by 0.03, which is 0.67: the statin group carries two thirds of the placebo risk.

Absolute risk reduction is 0.03 minus 0.02, one percentage point. Relative risk reduction is that 0.01 divided by the control risk of 0.03, giving 33%. The number needed to treat is 1 divided by 0.01, so 100 patients must take the statin for five years to prevent one additional heart attack.

Each number answers a different question. Relative risk asks how the groups compare proportionally. Absolute risk reduction asks how much an individual's probability actually changes. The NNT translates that into clinical workload: how many people must take the drug, with its costs and side effects, for one person to benefit. An evaluation that quotes only one of these numbers is incomplete.

How good is a given NNT? Values below 10 are considered highly effective: one person benefits for every ten treated. Values from 10 to 100 are moderately effective, worthwhile when the outcome prevented is serious. Above 100 the benefit is marginal for most patients, and cost and side effects dominate the decision. Whether an NNT is acceptable always depends on the severity of the outcome prevented.

Book notes
  • RR = risk in exposed group divided by risk in unexposed group. ARR = control risk minus treatment risk. RRR = ARR divided by control risk. NNT = 1 divided by ARR.
  • Tamoxifen P-1 trial: 13,388 women, 89 versus 175 cases, a 49% relative reduction but only 0.64 percentage points absolute, NNT 156.
  • Quote RR, ARR and NNT together; any one alone can mislead.

Fill the gap: [___] risk reduction equals control risk minus treatment risk, and it gives the real-world size of the difference.

Interactive · Relative Risk Calculator
5
Why relative risk misleads, and the numbers that keep you honest
analyse

Relative figures amplify small effects in low-risk populations. A supplement advertised as halving cancer risk sounds dramatic, but if the baseline risk is 2 in 100,000, halving it changes the risk to 1 in 100,000: an absolute reduction of one thousandth of a percentage point and an NNT of 100,000. The relative figure is truthful, yet stripped of the context that gives it meaning.

The related measure, attributable risk, is the difference in risk between exposed and unexposed groups: the excess risk that can be attributed to the exposure itself. If 20% of smokers and 2% of non-smokers develop a disease, the attributable risk is 18 percentage points. Public health agencies use it to estimate how much disease would disappear if the exposure were removed.

The 2002 Heart Protection Study followed 20,536 high-risk patients for five years. Simvastatin cut major vascular events from 25.2% to 19.8%, an absolute reduction of 5.4 percentage points and an NNT of about 19, so statins became standard care after a heart attack. In low-risk people the same relative reduction gives an NNT near 83, so prevention prescribing is more contested.

Do not confuse relative risk with relative risk reduction. RR is the ratio of the two risks; RRR is how far the risk falls, proportionally. A drug with an RR of 0.8 gives a 20% relative reduction, calculated as 1 minus 0.8, not an 80% reduction. Markers deduct for this slip every year.

A trial cuts a disease risk from 0.004% to 0.002%, a 50% relative risk reduction. Which evaluation is best?

6
Statistically significant is not the same as important
example

A p-value below 0.05 means there is less than a 5% probability of seeing a result this extreme if the treatment truly did nothing. It says nothing about whether the effect matters. With a large enough sample, even a trivially small difference becomes statistically significant, because random variation no longer hides it.

One trial of 500,000 patients found a new drug lowered average blood pressure by 0.3 mmHg more than placebo, with p = 0.001. The result is highly significant statistically and clinically meaningless: no patient feels or benefits from a 0.3 mmHg change. Statistical significance tells you an effect probably exists; clinical significance, judged from effect size, ARR and NNT, tells you whether it is worth acting on.

HSC exam move

When a stimulus quotes p < 0.05, never write "therefore the treatment works". Write that the result is unlikely to be due to chance, then evaluate the effect size and NNT before judging importance. Band 6 answers separate statistical significance from clinical significance every time.

True or false: a result with p < 0.05 is always clinically important.

7
From data to a fair claim
explain

A strong answer separates what the study found from what the study proves. "Associated with lower risk" is safer than "causes lower risk" unless the method controls competing explanations.

Common error A large sample removes every source of bias +

A large sample can reduce random sampling error, but it cannot repair systematic selection, measurement or recall bias.

Fix: judge sample size and bias separately, then explain how each affects confidence.
HSC exam move

For "evaluate the method", write: design, strength, limitation, judgement. Use exact evidence from the stimulus when it is provided.

8

Choose your route

8
Reading survival curves without being fooled
explain

A Kaplan-Meier survival curve shows the proportion of a study group that has not yet experienced the outcome, whether death, recurrence or hospitalisation, plotted against time. Every curve starts at 100% and steps downward as events occur. A line that falls more steeply means events are happening faster in that group, the worse outcome.

Three features carry the marks. The vertical gap between two lines at a time point is the difference in survival probability between the groups. A plateau means no observed events during that interval, often because follow-up ended, not because patients are cured. Small tick marks are censored patients who left the study or reached its end event-free; their later fate is unknown, not assumed to be survival.

In one melanoma trial, five-year survival was 52% on immunotherapy and 28% on chemotherapy, with the curves separating from six months onward. The honest conclusion is a 24 percentage point absolute difference in five-year survival, and the widening gap fits immunotherapy's mechanism of building a durable immune response. It does not show a cure: 48% of the immunotherapy group still died within five years.

Schematic Kaplan–Meier survival graph comparing immunotherapy and chemotherapy. Both curves start at 100 percent and step down when events occur. At five years they end at 52 and 28 percent, a 24 percentage-point gap. Short diagonal marks indicate censored participants whose later outcomes are unknown.
The endpoint values match the lesson example; intermediate steps are schematic. Read both groups at the same time, treat censor marks as unknown later outcomes, and never infer cure from a plateau alone.

Interrogate: At year five, state the absolute survival difference. Then point to one event step, one censor mark and one interval where a flat line would need the numbers at risk before interpretation.

Screening creates a subtler trap. A screening program that detects tumours earlier can lift five-year survival from 50% to 80% while mortality stays unchanged. Earlier diagnosis moves the diagnosis date forward without moving the date of death (lead-time bias), and some detected tumours would never have caused harm (overdiagnosis). Survival statistics improve; deaths do not.

Book notes
  • Kaplan-Meier curve: y-axis is proportion event-free, x-axis is time; steeper fall means a worse outcome.
  • Gap between lines is the survival difference; tick marks are censored patients whose fate is unknown.
  • Lead-time bias and overdiagnosis can lift survival statistics while mortality stays unchanged.

A screened group's five-year survival rises from 50% to 80%, but the mortality rate is unchanged. What best explains both observations?

9
From association to causation: Hill's criteria
classify

Association is not causation, so epidemiologists judge causal claims against Bradford Hill's criteria. The examinable ones include temporality (the exposure must precede the disease), strength of the association, a dose-response relationship, biological plausibility, coherence with existing knowledge, and consistency across independent studies. The more criteria an association satisfies, the stronger the case for causation.

Temporality and plausibility fail in a famous spoof: countries with higher chocolate consumption produce more Nobel laureates per capita. There is no plausible mechanism linking national chocolate intake to individual genius, and the correlation is confounded by wealth and education spending. The association is real in the data; the causal story is not.

Australia's AIHW reports higher rates of cardiovascular disease, diabetes and renal failure among Aboriginal and Torres Strait Islander peoples. Genetics contributes to individual risk, but the population-level disparity is driven mainly by social determinants: income, housing, education, racism and access to healthcare interacting with biology. Attributing the whole gap to genetics fails the plausibility and coherence tests and misdirects prevention.

Odd one out: three of these strengthen a causal claim under Hill's criteria. Click the one that does not.

10
The evaluation checklist that maps to marks
apply

Evaluating a study is not finding flaws for the sake of it; it is identifying what the study can and cannot establish. Work through a fixed checklist: is the design right for the question, is the sample large enough, is it representative, were participants and assessors blinded, is the control group appropriate, was follow-up long enough, and were confounders controlled?

Apply it to a real stimulus. A six-week RCT of 200 patients found a new anti-inflammatory drug cut self-reported knee pain 35% more than placebo (p = 0.03). The trial was single-blind and excluded patients with severe kidney disease. The strength earns the first mark: randomisation controls confounders, and an RCT is the right design for testing a treatment.

The limitations decide the band. Single-blind means the researchers knew the allocations and could rate a subjective outcome like pain more favourably, an assessment bias. Six weeks is short for a condition that often improves spontaneously. Excluding kidney patients limits generalisability. And p = 0.03 with 200 patients sits close enough to the threshold that sampling variation is a live concern.

Book notes
  • Checklist: design, sample size, representativeness, blinding, control group, follow-up length, confounding.
  • Single-blind trials risk assessment bias, especially for subjective outcomes like self-reported pain.
  • Exclusion criteria limit who the results generalise to.
HSC exam move

For "evaluate the method", structure the answer as design, two strengths with reasoning, two limitations with reasoning, then a judgement of what can and cannot be concluded. Use epidemiology vocabulary: confounding, bias, temporality, significance, representativeness. "The study was good" earns nothing.

Two truths and a lie: click the statement that is false.

Interactive · Study Type Classifier
11
Choose your route
differentiate

Pick one route, whichever matches how confident you feel right now. Supported gives you the most structure, Stretch asks for the most independent judgement. You only need to complete one.

Supported

Use the frame to evaluate one study method.

Cover Design: … Strength: … Limitation: … Judgement: …

Core

Compare a cohort study and an RCT for testing a prevention strategy.

Cover Cohort is useful because … RCT is stronger for … Limitation: …

Stretch

A headline reports a large relative risk reduction. Explain one extra number you would request before judging importance.

Cover I would request … because relative framing can …

9

Exit check

12
Exit check
retrieve
Memorise

Bias, confounder, control group, cohort, case-control, RCT.

Understand

Methods determine how strongly a study supports its conclusion.

Apply

Evaluate a method by naming strengths and limitations.

Avoid

Do not call a weak association proof of causation.

6

Independent practice

01
Multiple Choice
+5 XP

A fresh set drawn from this lesson's question bank, feedback shown immediately. +5 XP per correct · +25 XP all correct

Pick your answer, then rate your confidence, that tells the system what to drill next.

02
Short Answer, 15 marks
+5 XP

ApplyBand 4(4 marks) 1. A cohort study finds that an exposure is associated with disease. Use the evaluation frame: name the design, one strength, one limitation and a justified conclusion.

AnalyseBand 4–5(5 marks) 2. Compare a cohort study and an RCT for testing a prevention strategy. Include one strength and one limitation of each.

EvaluateBand 5–6(6 marks) 3. A headline claims a prevention program “cuts risk by 40%”. Explain one extra number you would request, then evaluate why a headline alone is not enough evidence for a health decision.

Show all answers

Multiple choice

MC answers and full explanations are shown inline as you complete each question. Use the retry button to attempt a fresh set from the lesson bank.

Short Answer Model Answers

SA1 (4 marks): The method is a cohort study [1]. Following exposed and unexposed groups over time can establish that exposure preceded disease [1]. Confounding or unequal loss to follow-up may partly explain the association [1]. The result supports an association, but the cohort method alone does not prove that the exposure caused the disease [1].

SA2 (5 marks): A cohort can follow naturally occurring exposure and establish temporal order, but confounding and attrition can weaken its conclusion [2]. An RCT uses random allocation to distribute confounders more evenly and is stronger for testing an intervention, but it may be unethical, impractical or less representative of the wider population [2]. The stronger method depends on the question and whether random allocation is ethical [1].

SA3 (6 marks): Request the baseline risk or absolute risk reduction, because a 40% relative reduction may describe a small absolute change [1]. Check the study design, comparison group, sample, follow-up, outcome measure and whether the result is practically important [2]. Bias, confounding or selective reporting may weaken the headline [1]. A health decision also needs evidence about harms and whether the participants represent the target population [1]. Therefore the headline is not enough; the full method and absolute effect are needed before judging the claim [1].

7

Retrieve and reflect

Check what actually stuck
Take the full module quiz
quiz

A full module quiz covering every lesson in this module, not just this one. Set aside a decent block of time and treat it like a real assessment.

Start the module quiz →
Blast the Correct Answer
blaster

Defend your ship by identifying study designs, strengths, limitations and fair conclusions. Scores count toward the Asteroid Blaster leaderboard.

☄️ Play Asteroid Blaster →
Race Through Study Evaluation!

Answer questions on cohort, case-control and randomised studies, then evaluate bias and confounding. Pool: lessons 1–13.

How did your thinking change?

Return to the breakfast survey from Think First. It reported an association between eating breakfast and lower heart-disease rates, but breakfast habits may also be linked with exercise, income and healthcare access.

Evaluate the method using the four-part frame from this lesson: name the likely design, state one strength, identify one limitation or confounder, and finish with a justified conclusion that does not overclaim causation.