Austin Bradford Hill gave epidemiology nine famous ways to look at an association. Then, on the same page, he warned readers not to turn them into rules.
That second sentence is the one teaching summaries often lose. Hill's 1965 address, The Environment and Disease: Association or Causation?, is routinely compressed into the “Bradford Hill criteria”: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. The original paper calls them viewpoints. None, Hill says, can settle causation on its own; none should operate as a required box in a scorecard.[1][3]
Read from beginning to end, the address is doing something more useful than handing out a verdict machine. It moves through three different jobs. First, decide whether an observed pattern deserves a causal interpretation. Second, look hard for another explanation—chance, selection, measurement, or some feature travelling with the exposure. Third, decide whether the present evidence is strong enough to act, given what happens if action is right or wrong. Later epidemiologists sharpened and formalized those jobs, but they also noticed how easily Hill's flexible prompts had hardened into a ritual checklist.[3][4][5]
A speech after the smoking verdict, not before it
Hill delivered the address to the Royal Society of Medicine's new Section of Occupational Medicine on 14 January 1965; it appeared in print that May.[1] The timing matters. On 11 January 1964, U.S. Surgeon General Luther Terry had released Smoking and Health. Its advisory committee reviewed more than 7,000 articles and concluded, among other things, that cigarette smoking caused lung cancer in men.[2] Hill was not inventing nine gates so that smoking could finally pass through them. He was reflecting on how investigators had moved from repeated observations to a causal judgment—and on how preventive medicine should make similar judgments about less settled hazards.
The paper therefore opens with a practical question. If an undesirable event B appears more often in the presence of environmental feature A, would changing A alter the frequency of B? Hill does not require the full biological chain to be known before that question can be answered or before prevention begins. Sometimes mechanism is clear; sometimes observations have to carry more of the argument.[1]
This framing places the paper between two dates in the history of causal inference. The 1964 smoking report supplied a recent, public example of evidence synthesis. The 1965 address generalized lessons from it. By 2004–2005, commentators were explicitly pushing back against the afterlife of the “criteria,” arguing that causal inference is better understood through study validity, competing explanations, and estimation than through a criterion count.[3][4]
Strength is a clue, and the scale changes the story
Hill begins with strength of association, but his own examples immediately complicate it. In the British doctors' evidence he cites, lung-cancer mortality among cigarette smokers was roughly 9 to 10 times the rate among nonsmokers, and among heavy smokers roughly 20 to 30 times as high. Coronary-thrombosis mortality, by contrast, was no more than about twice as high in smokers. A large ratio makes it harder to imagine that an unmeasured companion of smoking explains the whole lung-cancer excess. It does not prove that no such companion exists.[1]
Then Hill changes the denominator. If the question is not etiology but how many additional deaths smoking produces, the absolute rates matter. He gives annual lung-cancer death rates per 1,000 doctors of 0.07 among nonsmokers, 0.57 among those smoking 1–14 cigarettes a day, 1.39 among those smoking 15–24, and 2.27 among those smoking 25 or more. The ratios are highly informative about causal interpretation; the absolute rates make the burden difference recoverable. Neither scale is universally “best.”[1]
John Snow's cholera evidence makes the same point with an even sharper contrast: 71 deaths per 10,000 houses supplied by one water company versus 5 per 10,000 among houses supplied by another. Yet Hill turns the coin over. A weak association can still be causal. Only a small share of people carrying meningococci develop meningitis, and only a small share of people exposed to rat urine contract leptospirosis. Dismissing a hazard because the observed association is slight would turn “strength” from a clue into an exclusion rule Hill never endorsed.[1]
Consistency receives the same treatment. Hill notes that the smoking–lung-cancer association had appeared in 29 retrospective and 7 prospective inquiries. Repetition across people, places, circumstances, and methods weakens the case for one persistent error. But identical results from studies that repeat the same bias are not independent confirmation, while a sound result need not disappear merely because a differently selected comparison group produces another answer.[1]
The middle of the paper dismantles one-cause, one-disease thinking
Specificity sounds strongest when one exposure points to one disease in one group of workers. Hill's chimney-sweep and nickel-refinery examples use that geometry. But he promptly explains why it cannot be demanded. One agent can cause several outcomes; one disease can have several causes. Milk can carry different infections. Smoking can affect more than the lung. Scrotal cancer can occur without chimney soot. The absence of a one-to-one relationship is not an acquittal.[1]
Temporality asks which came first. Did an occupation promote tuberculosis, or did people already vulnerable to tuberculosis select that work? Did a diet precede disease, or did early disease change diet? Hill treats these as questions that may be difficult, especially with slow disease. Modern causal-inference writing goes one step further: of his nine viewpoints, temporality is the one logically necessary condition, because a cause must precede its effect.[4][5]
Biological gradient asks whether more exposure tracks with more disease. Plausibility asks whether current biology can accommodate the proposed pathway. Coherence asks whether the causal interpretation conflicts with the disease's known history and biology. Experiment asks whether removing or changing the exposure alters outcomes. Analogy asks whether a similar exposure–outcome relationship makes the new claim less surprising.[1]
These are useful angles of inspection, but the paper repeatedly shows why absence is not disproof. A genuine causal relation may have a threshold, plateau, or other non-monotonic pattern, while broad exposure categories can hide that pattern. A mechanism can look implausible only because biology has not caught up: Snow's waterborne account of cholera preceded isolation of the responsible organism by decades. Human evidence for arsenic and skin cancer did not become void because an animal model was missing. Analogy can suggest where to look, but resemblance cannot establish that two causal systems actually work alike.[1][4][5]
The connective tissue is a search for rival explanations
The nine labels occupy the memory of the paper. The examples between them contain its method.
When Hill discusses consistency, he describes a hospital comparison that could mislead even if repeated precisely. Patients admitted for ulcer surgery and patients admitted for uncomplicated hernia surgery were asked about recent emotional crises. A person in crisis might still seek urgent treatment for an ulcer but postpone an elective hernia repair. The comparison groups would then differ because of the route into hospital, not because emotional crisis caused ulcer. Repetition would reproduce the selection problem rather than cure it.[1]
When he discusses significance testing, Hill draws another boundary. A formal test can estimate how compatible data are with a chance model and help describe precision. It cannot certify that the study measured the right people, the right exposure, or the right outcome. Hill recalls an occupational study in which cardroom workers had more than three times as much respiratory sickness absence as less-exposed workers in the same mills, while non-respiratory absence was similar. The pattern was specific enough that a significance test would have added little to the argument. Elsewhere, he warns that a small p value cannot repair an unknown volunteer fraction, the 20% of patients lost to follow-up, or the 30% of a random sample who were never contacted.[1]
This is where the original address comes closest to modern causal reasoning. Before interpreting an association, ask whether selection, confounding, measurement, or chance could have produced it. Define the causal question clearly. Establish the time order. Measure the effect on a scale suited to that question. Make assumptions visible. Hill's viewpoints can organize some of that scrutiny, but they cannot replace it.[4][5]
The checklist refusal is explicit
After the ninth viewpoint, Hill pauses to say exactly what his list is not. It cannot provide indisputable evidence for or against cause and effect. He rejects “hard-and-fast rules” and says no viewpoint is a sine qua non. The list is meant to help a reader ask whether another explanation fits the facts equally well or better.[1]
That wording leaves one correctable weakness. Later authors point out that temporality really is indispensable in a causal claim, even if the other eight are neither necessary nor sufficient.[4][5] The correction does not turn the other viewpoints into a scoring system. It clarifies the difference between a logical requirement and evidentiary clues whose value changes with context.
The distinction matters because checklists invite arithmetic. Six “criteria” present may sound stronger than four. Hill provides no basis for that sum. A biologically plausible story can be wrong; the same flawed design can produce consistent estimates; a strong association can arise from severe bias; a real cause may lack specificity or a simple gradient. Evidence has structure, not merely quantity.[3][4]
The neglected final move is from inference to action
Hill's last section changes the question. Scientific appraisal, he argues, should judge evidence on its merits, independent of who benefits from the conclusion. A practical decision must also consider the consequences of being wrong.[1]
That is why different interventions can justifiably demand different evidentiary thresholds. Hill's own, now-dated example was that relatively slight evidence might justify restricting a morning-sickness drug because he assumed little harm if the causal inference proved wrong—an assumption later commentators specifically questioned. He said fair evidence might justify replacing a probably carcinogenic workplace oil, but very strong evidence would be needed before requiring people to burn a household fuel they disliked or abandon cigarettes and favored foods. The threshold is not a property of causation itself; it belongs to the decision built on top of the causal judgment.[1][3]
This is not permission to make science say whatever policy requires. It is a demand to keep two ledgers. One records how credible the causal explanation is. The other records the expected harms of acting and waiting. Uncertainty appears in both. Hill's closing point is that incomplete science does not create a right to ignore what is already known or to postpone action that the available evidence makes reasonable.[1]
The durable reading of the paper is therefore not “check nine boxes.” It is: state the causal question; test the study's validity; establish time order; compare the causal account with serious rivals; use pattern, mechanism, and intervention evidence in context; then make the action threshold explicit. Hill's nine viewpoints still earn their place—as windows, not turnstiles.
Sources
- Austin Bradford Hill, “The Environment and Disease: Association or Causation?” Proceedings of the Royal Society of Medicine, 1965 — Columbia University scan of the complete primary address, including the nine viewpoints, significance-testing discussion, and case for action.
- Centers for Disease Control and Prevention, “A History of the Surgeon General's Reports on Smoking and Health” — official account of the 11 January 1964 report, its evidence base, and its causal conclusions.
- Carl V. Phillips and Karen J. Goodman, “The missed lessons of Sir Austin Bradford Hill,” Epidemiologic Perspectives & Innovations, 2004 — analysis of the checklist misreading, significance, validity, and action under uncertainty.
- Kenneth J. Rothman and Sander Greenland, “Causation and causal inference in epidemiology,” American Journal of Public Health, 2005 — modern account of multicausality and why causal inference is not a criterion-counting process.
- Sanghyuk Bae et al., “Causal inference in environmental epidemiology,” Environmental Health and Toxicology, 2017 — contemporary treatment of study validity, Hill's viewpoints, temporality, bias, and general versus individual causation.
- Wellcome Library, “Sir Austin Bradford Hill, portrait,” Wikimedia Commons — source page for the archival photograph used as the article image.