Year 12 Biology Module 8 · IQ3 ⏱ ~45 min Practice bank · 3 Short Answer Lesson 12 of 21 Data & methods focus

Epidemiology: Measuring Disease Patterns

Epidemiology uses population data to measure how common disease is, who is affected and whether patterns are changing. Learn the core measures before judging causes or treatments.

Today's hook: If a headline says "more people have diabetes", what else do you need to know before deciding whether risk has actually increased?
0/5TASKS
1
You’re here

Get oriented

Warm up first

Three quick questions from earlier lessons. Pulling old material back to mind before you learn something new makes the new material stick better, so this is not busywork.

Worksheets

Practise this lesson

Four printable worksheets that build from the foundations up to exam-style questions, start at whatever level suits you.

Lesson map

Count -> compare -> question

First learn the three measures. Then use them to interpret patterns and avoid weak conclusions.

  1. Choose the measure.Incidence, prevalence and mortality answer different questions.
  2. Compare fairly.Rates and age-standardisation matter more than raw totals.
  3. Check the evidence.Study design affects what conclusions are justified.

Know what matters

Must Know
  • Incidence counts new cases over a time period.
  • Prevalence counts existing cases at a point or across a period.
  • Mortality measures deaths from a disease.
  • Population patterns need rates, not only raw totals.
Should Know
  • Age-standardised rates allow fairer comparisons between populations.
  • Prevalence can rise when treatment helps people live longer.
  • Study design affects whether evidence supports association or causation.
Going Deeper
  • Bradford Hill criteria as a causation framework.
  • Relative risk and confounding in observational studies.
  • Why RCTs are not ethical for harmful exposures.
0
Predict first: which number matters?
connect

A town grows from 10,000 people to 20,000 people. Cancer cases also double. What should you check before saying cancer risk increased?

1
Think first: is "more diabetes" the same as "diabetes is getting worse"?
connect

Between 2000 and 2022 the total number of Australians diagnosed with Type 2 diabetes more than doubled, and headlines reported this as a worsening epidemic. Over the same period Australia's population grew and aged, and far more people were screened and diagnosed than ever before.

A researcher argues that once you adjust for population size and age structure, the age-standardised incidence of Type 2 diabetes has been roughly stable, and even declining in some age groups. Hold both claims in mind as you work through this lesson, because they can both be true at once.

Q1: What is the difference between the total number of cases of a disease and the rate of disease in a population? Why does this distinction matter for public health decisions?

Q2: Why might improved screening and diagnosis make a disease appear to be increasing even if the underlying rate is unchanged?

5

Key vocabulary, translated

2
Key vocabulary, translated
vocab
EpidemiologyThe study of who gets a disease, where, when and why, across whole populations rather than one patient. It hunts for patterns in numbers.Like this: comparing lung cancer rates in smokers and non-smokers across thousands of people is epidemiology. Treating one patient's tumour is medicine.
IncidenceThe number of NEW cases arising in a population during a set time. It shows how fast a disease is appearing.Like this: the number of children per 100,000 newly diagnosed with type 1 diabetes in a year is an incidence figure. Prevalence would also count everyone already living with it.
PrevalenceThe total number of existing cases at one point in time, new and long-standing together. It shows how much disease the population is carrying.Like this: a lifelong disease builds high prevalence even when incidence is low, because cases accumulate and nobody leaves the count.
MortalityThe number of deaths from the disease in a population over a period. High incidence with low mortality means common but survivable.Like this: melanoma incidence in Australia is high while mortality is far lower, because most cases are removed early.
ConfounderA third factor linked to both the exposure and the disease, which can invent an apparent relationship or hide a real one.Like this: people who drink more coffee also tend to smoke more, so smoking can make coffee look like a cause of lung cancer when it is not.

True or false: prevalence can increase even if the number of new cases each year falls.

3
Match the measure to the question
apply

Incidence

Use when asking, "How many new cases are appearing?"

Prevalence

Use when asking, "How many people currently live with it?"

Mortality

Use when asking, "How many people die from it?"

Remember!

A raw count describes how many events occurred; a rate relates those events to the population at risk.

Pause and copy the highlighted distinction before interpreting a trend.

Build a data answer+7 XP

Put the data-interpretation steps in a useful HSC order.

  • Explain what conclusion is supported and what is not proven.
  • Identify the measure used.
  • Quote values and describe the pattern.
  • Check whether the data are totals, rates or age-standardised rates.
4
Incidence and mortality: counting new cases and deaths
explain

When Doll and Hill published their 1950 case-control study linking smoking to lung cancer, they had a statistical association but no proof of causation. Generating any causal evidence began with measuring the disease correctly: counting new cases, existing cases and deaths in a defined population. Without those three measurements, no rate, no fair comparison and no causal claim is possible.

Incidence, new cases over time

Incidence is the rate at which new cases arise in a population over a defined time period, and it answers how fast a disease is developing. The formula is new cases in the period, divided by the population at risk, multiplied by 100,000. If 1,800 people are newly diagnosed with melanoma in a population of 10 million in one year, the incidence rate is 18 per 100,000 per year.

Use incidence when measuring the risk of developing a disease, tracking whether it is becoming more or less common, or evaluating a prevention program. One caution matters in exams: rising incidence can reflect genuinely increasing disease, or improved screening and diagnosis detecting cases that would previously have been missed.

Mortality rate, deaths from disease

Mortality rate counts deaths attributable to a specific disease per unit of population per unit of time: deaths from the disease, divided by population, multiplied by 100,000 per year. If 1,800 people die of coronary heart disease in a population of 10 million, the mortality rate is 18 per 100,000 per year. It measures severity and tracks whether treatment advances are saving lives.

Incidence and mortality can move independently. Most skin cancers have high incidence but low mortality, because they are common yet rarely fatal if caught early. Pancreatic cancer shows the reverse pattern, low incidence but around 90 percent mortality within five years. Quoting that contrast is a fast way to show a marker you understand what each measure actually captures.

Book notes
  • Incidence = new cases in a period, divided by population at risk, times 100,000 per year.
  • Mortality rate = deaths from the disease, divided by population, times 100,000 per year.
  • High incidence with low mortality: most skin cancers. Low incidence with high mortality: pancreatic cancer.
  • Rising incidence can reflect better screening, not only more disease.

A disease has low incidence but very high mortality. Which statement is consistent with this pattern?

5
Prevalence, and why better treatment can push it up
explain

Prevalence is the total proportion of a population living with a condition at a given time, including both new and existing cases. The formula is existing cases divided by total population, multiplied by 100. If 1.3 million of Australia's roughly 26 million people have Type 2 diabetes at a given time, prevalence is 5 percent. It answers how much disease exists in the community right now.

Prevalence is the planning measure. Health departments use it to decide how many people need insulin, dialysis, cancer treatment or aged care, and to allocate resources across the system. Incidence tells you about risk for an individual; prevalence tells you about burden on the health system. Exam answers improve the moment you state which question a measure actually answers.

Prevalence is approximately incidence times duration

The key relationship is that prevalence is approximately incidence multiplied by average disease duration. A disease with low incidence but long duration, such as Type 2 diabetes, accumulates a large pool of existing cases and reaches high prevalence. A disease with high incidence but short duration, such as influenza, resolves or kills quickly and never builds the same pool.

This explains a pattern students find counterintuitive. Effective treatment that extends life raises prevalence even when incidence is stable or falling, because patients remain in the existing-cases pool for longer. HIV in high-income countries is the classic example: antiretroviral therapy means fewer people die, so prevalence rose through the 2000s while new infections fell.

Book notes
  • Prevalence = existing cases, divided by total population, times 100 (a percentage).
  • Prevalence is approximately incidence multiplied by average duration of disease.
  • Effective treatment extends survival, so prevalence can rise while incidence falls (HIV on antiretroviral therapy).
  • Use prevalence for health-system planning; use incidence for risk and prevention.

Odd one out: three of these statements about prevalence are correct. Click the one that is not.

6
Age-standardisation: comparing populations fairly
example

Raw, or crude, rates cannot always be fairly compared between populations with different age structures. An older population will show higher crude rates of cancer, cardiovascular disease and dementia simply because it contains more older people, not because disease risk at any given age is higher. Age-standardisation applies a standard age distribution to both populations so the underlying rates can be compared on a level playing field.

Australia illustrates the difference clearly. The total number of cancer deaths has risen for decades because the population is larger and older, yet the age-standardised cancer mortality rate has been falling over the same period, because improved prevention, screening and treatment have reduced the death rate per case. Both statements are true; they answer different questions.

When an exam table gives both crude and age-standardised columns, treat the gap between them as information. A large gap tells you the populations being compared have different age structures, and the age-standardised column is the one on which a fair conclusion must rest.

Book notes
  • Crude rate: events divided by total population, no adjustment.
  • Age-standardised rate: adjusted to a standard age distribution, so populations compare fairly.
  • Australia: total cancer deaths rising, age-standardised cancer mortality falling, both true at once.

Fill the gap: to compare disease rates fairly between populations with different age structures, epidemiologists use [___] rates.

Interactive · Epidemiology Calculator
7
Worked example: reading an epidemiological table
analyse

HSC exams regularly hand you a table of epidemiological data and ask you to describe, analyse or evaluate it. The table below shows Type 2 diabetes figures for Australia across two decades, and the disciplined habit is to read every column before describing any trend.

YearTotal diagnosed casesPopulation (millions)Crude prevalence (%)Age-standardised prevalence (%)
2000640,00019.23.3%4.1%
2010970,00022.34.4%4.3%
20221,300,00025.95.0%4.2%

Total diagnosed cases roughly doubled, from 640,000 to 1.3 million, but the population also grew from 19.2 to 25.9 million, so part of the rise is population growth alone. Crude prevalence rose from 3.3 to 5.0 percent, partly because the population aged and older Australians have higher Type 2 diabetes rates.

The age-standardised column tells the quieter story: prevalence adjusted for age structure moved only from 4.1 to 4.2 percent. Much of the apparent epidemic reflects demographic change rather than a dramatic rise in underlying risk at any given age. That single comparison is often the difference between a Band 4 and a Band 6 interpretation.

HSC exam move

When asked to analyse epidemiological data: describe the overall trend, quote specific values, compare crude with age-standardised columns, then state what the data can and cannot support, naming confounders, age structure and correlation versus causation as limits.

In the table, total cases roughly doubled but age-standardised prevalence barely changed. The best conclusion is:

7

Study design changes the strength of the claim

8
Study design changes the strength of the claim
explain

Observational studies can reveal patterns and associations in real populations. Randomised controlled trials are stronger for testing treatments, but they cannot be used when assigning a harmful exposure would be unethical.

Common error An association proves that the exposure caused the disease +

A third factor may influence both the exposure and the outcome.

Fix: name a plausible confounder and limit the conclusion to an association unless the method supports causation.
HSC exam move

When evaluating epidemiological evidence, state the design, identify a strength, identify a limitation, and judge whether the conclusion is causal or only associative.

9
Four study designs, from observation to experiment
explain

Epidemiologists cannot randomly assign people to smoke cigarettes or eat poorly for decades, so most questions about disease and exposure must be studied observationally. The design a researcher chooses determines which questions can be answered, and how strong a conclusion the resulting evidence can support.

Cohort study, prospective

A group of disease-free people is followed forward in time, and exposed and unexposed subgroups are compared for disease development. This establishes temporal sequence, exposure before disease, and suits common outcomes. The costs are time and money, plus loss to follow-up, and rare diseases are impractical. The British Doctors Study followed about 40,000 doctors from 1951, comparing smoking status with lung cancer rates over decades.

Case-control study, retrospective

People with a disease, the cases, are compared with disease-free controls, and past exposures are compared between the groups. The design is quick, inexpensive and efficient for rare diseases, and several exposures can be studied at once. Its weaknesses are recall bias, because cases remember exposures differently from controls, and the inability to calculate incidence directly. Comparing asbestos exposure history in mesothelioma patients with matched controls is the textbook example.

Cross-sectional study

Exposure and disease are measured at the same point in time, a snapshot of the population. It is fast and cheap, measures prevalence well and generates hypotheses for further study, but it cannot show which came first, so it cannot establish temporal sequence or calculate incidence. Australia's National Health Survey, recording smoking status and cardiovascular disease in one sample at one time, works this way.

Randomised controlled trial

Participants are randomly allocated to intervention or control groups, and randomisation distributes known and unknown confounders evenly by chance, which is why the RCT is the gold standard for establishing causation. Double-blinding reduces bias further. The limits are ethics, cost and generalisability: no ethics board would approve assigning people to a harmful exposure. HPV vaccine trials, randomising participants to vaccine or placebo and comparing precancerous lesion rates, show the design at full strength.

Book notes
  • Cohort: follow exposed and unexposed forward; establishes temporal sequence; slow and costly.
  • Case-control: compare past exposure in cases versus controls; efficient for rare disease; recall bias.
  • Cross-sectional: one-time snapshot; measures prevalence; cannot show which came first.
  • RCT: randomisation controls confounders; gold standard; unethical for harmful exposures.

Why can a randomised controlled trial not be used to test whether smoking causes lung cancer?

Beyond the syllabus. The Bradford Hill criteria, the prevalence–incidence–duration approximation and the detailed bias taxonomy are extension. For the exam you need incidence, prevalence and mortality, and the broad strengths and limitations of epidemiological study designs.
10
Confounding, bias and the limits of association
analyse

A confounding variable is associated with both the exposure and the outcome, and it can create a spurious or distorted association. The classic case: studies once found coffee drinking associated with lung cancer. Coffee drinkers of that era were also far more likely to smoke, and smoking is causally linked to lung cancer. Once smoking is controlled for, the coffee association largely disappears.

Confounders can be controlled by matching cases and controls on the confounding variable, by statistical adjustment, by stratified analysis, or, most powerfully, by randomisation, which distributes confounders equally between groups by chance. In an exam, naming a plausible confounder and the method that would control it earns credit that naming the association alone does not.

Four biases worth naming

  • Selection bias: the sample does not represent the target population; the healthy worker effect leads occupational studies to underestimate disease rates.
  • Recall bias: cases remember past exposures differently from controls, a particular problem in case-control studies.
  • Information bias: systematic errors in measuring exposure or outcome, misclassifying disease status or exposure level.
  • Reporting bias: positive results are more likely to be published than null findings, distorting the visible evidence base.

Correlation is not causation

Two variables can be statistically associated without one causing the other. Ice cream sales correlate with drowning deaths because both rise in hot weather, a confounder. Establishing causation needs the Bradford Hill criteria: strength of association, consistency across studies, temporality, dose-response and biological plausibility. Tobacco meets them all, with smokers at 15 to 25 times the lung cancer risk of non-smokers, a clear dose-response with pack-years, and carcinogens in smoke that form DNA adducts and TP53 mutations.

Book notes
  • Confounder: associated with both exposure and outcome (smoking confounds coffee and lung cancer).
  • Control methods: matching, statistical adjustment, stratification, randomisation.
  • Biases: selection, recall, information, reporting.
  • Causation needs Bradford Hill criteria: strength, consistency, temporality, dose-response, plausibility.

Fill the gap: a variable associated with both the exposure and the disease outcome, which can create a false apparent association, is called a [___] variable.

11
Doll and Hill: the cohort study that changed public health
example

In 1951, Richard Doll and Austin Bradford Hill sent questionnaires about smoking habits to every doctor on the British Medical Register, then followed about 40,000 of them for decades, recording causes of death. It was one of the first large prospective cohort studies, and it produced the most compelling epidemiological evidence for the smoking and lung cancer causal link.

Within four years the data were clear enough that Doll, himself a smoker, quit. After 50 years of follow-up, the study had quantified that smoking cuts life expectancy by about 10 years, established the dose-response between pack-years and lung cancer mortality, with a relative risk of 14.9 in heavy smokers, and showed that quitting before age 35 restores near-normal life expectancy.

The design did the causal work. Following people forward in time established that smoking preceded the cancer, ruling out reverse causation, and a well-defined professional cohort with reliable death certification minimised selection and information bias. Consistency across subgroups, a clear dose-response and a known mechanism satisfied the Bradford Hill criteria without any randomised trial, which would have been unethical anyway.

True or false: the smoking and lung cancer causal link was established by randomly assigning doctors to smoke or not smoke.

Interactive · Study Design Matcher
8

Choose your route

12
Choose your route
differentiate

Pick one route, whichever matches how confident you feel right now. Supported gives you the most structure, Stretch asks for the most independent judgement. You only need to complete one.

Supported

Choose incidence, prevalence or mortality for each question.

Cover New cases = … Existing cases = … Deaths = …

Core

Explain why raw totals can mislead when populations differ.

Cover Raw totals can mislead because … A fairer comparison uses …

Stretch

A study finds exposure X is associated with disease Y. Evaluate one reason this may not prove causation.

Cover The association may be affected by … This limits the conclusion because …

9

Exit check

13
Exit check
retrieve
Memorise

Incidence, prevalence, mortality, confounder.

Understand

Different measures answer different population questions.

Apply

Interpret a trend using rates and quoted data values.

Avoid

Do not infer causation from association without evaluating the study.

6

Independent practice

01
Multiple Choice
+5 XP

A fresh set drawn from this lesson's question bank, feedback shown immediately. +5 XP per correct · +25 XP all correct

Pick your answer, then rate your confidence, that tells the system what to drill next.

02
Short Answer, 14 marks
+5 XP

ApplyBand 4(4 marks) 1. Distinguish incidence, prevalence and mortality. Then explain why prevalence can rise even when incidence is falling.

AnalyseBand 4–5(5 marks) 2. A researcher is investigating whether regular physical activity reduces the risk of Type 2 diabetes. Describe how you would design a cohort study to investigate this question. Identify the cohort, the exposure and outcome variables, how data would be collected, and what would constitute evidence of an association. Identify one confounding variable and explain how it would be controlled.

EvaluateBand 5–6(5 marks) 3. Evaluate the following claim using your knowledge of epidemiological evidence and study design: "Because an RCT is the gold standard for medical evidence, we should require RCT evidence before accepting any claim that an environmental exposure causes disease."

Show all answers

Multiple choice

MC answers and full explanations are shown inline as you complete each question. Use the retry button to attempt a fresh set from the lesson bank.

Short Answer Model Answers

SA1 (4 marks): Incidence is the rate of new cases arising in a defined population over a specified time, (new cases ÷ population at risk) × 100,000, it measures how fast disease develops. Prevalence is the total proportion with the disease at a given time, (existing cases ÷ total population) × 100, it measures how much disease exists [2]. Why effective treatment raises prevalence despite falling incidence: prevalence ≈ incidence × duration. Effective treatment extends survival, so patients remain in the existing-cases pool for longer; even if incidence falls, the pool grows [1]. Example: HIV in high-income countries, antiretroviral therapy extended life, so prevalence rose through the 2000s while incidence (new infections) fell. The same pattern occurs for Type 2 diabetes (better treatment → longer survival → rising prevalence despite stable incidence) [1].

SA2 (5 marks): Cohort: recruit a large sample (50,000+) of adults aged 35–65 without T2D, willing to be followed 15–20 years [1]. Exposure: measure physical activity at baseline and every ~2 years (questionnaires or accelerometers), type, duration, intensity, frequency; classify into activity categories [1]. Outcome: development of T2D (fasting glucose ≥7.0 mmol/L, HbA1c ≥48 mmol/mol, or diagnosis), measured at each follow-up [1]. Evidence of association: compare annual T2D incidence in high- vs low-activity groups; calculate relative risk (<1.0 supports protection); test dose-response [1]. Confounding variable: diet (healthier eaters exercise more AND have lower T2D risk). Control: collect dietary data and statistically adjust, or restrict analysis to similar dietary patterns [1].

SA3 (5 marks): RCTs are the gold standard because randomisation distributes known and unknown confounders equally by chance, and blinding prevents bias, establishing causation [1]. But RCTs cannot ethically be used for harmful exposures: you cannot assign people to smoke, inhale asbestos, or receive high UV exposure for decades; an ethics board would never approve it. Requiring RCT evidence would mean we could never establish causation for environmental carcinogens experimentally [2]. Observational evidence can establish causation via the Bradford Hill criteria, strength, consistency, temporality, dose-response, biological plausibility, specificity. The smoking–lung cancer link was established entirely through observational cohort studies (Doll and Hill) plus mechanistic evidence, with no RCT [1]. Conclusion: the claim is partly valid (RCTs are ideal when ethical, drugs, vaccines, interventions) but inappropriate as a universal standard for harmful exposures; the appropriate standard is convergent evidence from multiple study types satisfying the Bradford Hill criteria [1].

7

Retrieve and reflect

Check what actually stuck
Take the full module quiz
quiz

A full module quiz covering every lesson in this module, not just this one. Set aside a decent block of time and treat it like a real assessment.

Start the module quiz →
Race Through Epidemiology!

Sprint through questions on incidence, prevalence, mortality and study design. Pool: lessons 1–12.

How did your thinking change?

Return to your Think First responses and consider the Bradford Hill 1965 framework in context. The Doll and Hill British Doctors Study generated a relative risk of 14.9 for lung cancer in heavy smokers, a finding that easily met Bradford Hill's criteria for strength, dose-response, temporality, and biological plausibility. Without the epidemiological measurement tools in this lesson (incidence, prevalence, RR, confounders), that landmark finding could never have been generated or evaluated.

  • Q1, total cases vs rate: The Doll/Hill study controlled for population size using rates (cases per 100,000 doctor-years). Total case count is influenced by population size, rate controls for this, allowing valid comparison across different populations and over time.
  • Q2, improved screening making disease appear to increase: Screening detects cases that previously existed but were undiagnosed. When screening uptake increases, the diagnosed (recorded) prevalence rises even if true prevalence is stable, this is ascertainment bias (a confounding variable Bradford Hill's criteria require you to rule out).
  • Write the formulas for incidence rate and prevalence from memory, and state in one sentence why age-standardised rates are more useful than crude rates for comparing the Doll/Hill 1950 cohort (older male doctors) to a modern mixed-age general population.