Sign in to save your progress and let facilitators follow along.Sign in
Week 2Lesson 10 of 54
Reading

Useful concepts

Start reading

The Singer chapter introduces the idea of measuring impact. Now we need the technical vocabulary. You’ll encounter these concepts repeatedly—in the GiveWell model you’ll trace shortly, in the quiz, and in your exercises.

Theory of change

A theory of change makes explicit how an intervention is supposed to work. It answers: what will you do, why do you believe it leads to impact, and what assumptions are you making?

Malaria Consortium’s theory of change for seasonal malaria chemoprevention (SMC): “If we distribute preventive antimalarial drugs to children during peak transmission season, and if caregivers administer them correctly, then fewer children will contract malaria, leading to fewer deaths.”

Each “if” is an assumption that could fail. Children might not take the medication. The drugs might be counterfeit. Malaria transmission patterns might shift. A good theory of change surfaces these assumptions so they can be tested.

The evidence hierarchy

Not all evidence is equally trustworthy. Researchers think about evidence quality as a hierarchy:

  • Strongest: Systematic reviews combining multiple rigorous studies; large, well-conducted randomised controlled trials (RCTs), especially those replicated across contexts.

  • Moderate: Individual RCTs; quasi-experimental designs (natural experiments, before-and-after with comparison groups); well-designed observational studies.

  • Weaker: Before-and-after comparisons without control groups; correlational studies; case studies.

  • Weakest: Theory and plausibility alone; testimonials and cherry-picked success stories.

When assessing impact claims, ask: where does the evidence sit on this hierarchy?

Why RCTs matter

An RCT solves the fundamental problem of causal inference. Suppose participants in a job training programme earn more money afterwards. Did the programme cause this? Maybe—but maybe those people would have found better jobs anyway.

By randomly assigning people to treatment and control groups, an RCT ensures the groups are similar in every way except the intervention. Any difference in outcomes can therefore be attributed to the intervention itself.

RCTs aren’t always possible (ethical constraints, practical limitations, cost) and they don’t tell you whether an intervention is the best use of resources. But when RCT evidence exists, we should weigh it heavily.

What outcomes might we measure?

QALYs (Quality-Adjusted Life Years) measure years of life adjusted for health quality. One QALY equals one year in perfect health; a year lived with a condition that reduces quality of life by 30% counts as 0.7 QALYs. Health systems use QALYs to decide which treatments to fund—the UK’s NHS, for example, typically funds interventions costing less than £20,000–30,000 per QALY gained.

DALYs (Disability-Adjusted Life Years) measure years of healthy life lost to disease or disability. They’re the mirror image of QALYs: where QALYs count health gained, DALYs count health lost. Global health organisations like the WHO use DALYs to compare disease burden across populations. GiveWell doesn’t use DALYs directly — they’ve developed their own moral-weights framework (which we’ll come to in Lesson 4) — but the idea is similar in spirit: translate different kinds of health outcomes into a common unit so they can be compared.

WELLBYs (Wellbeing-Adjusted Life Years) measure subjective wellbeing directly—one WELLBY equals one year at maximum life satisfaction. This allows comparisons across interventions that affect welfare but not health, like treating depression or reducing loneliness.

Each metric encodes assumptions. QALYs and DALYs assume health states can be quantified and that disabilities reduce welfare predictably—but disability advocates point out that people with disabilities often report higher wellbeing than these metrics assume. WELLBYs avoid this by measuring subjective experience directly, but assume self-reports are valid and comparable across people. None of these metrics is neutral.