Test characteristics
Build the 2 × 2 table with disease across the top and test result down the side.
- Sensitivity = TP / (TP + FN). Of those with disease, how many test positive. A highly SeNsitive test, when Negative, rules disease out — SnNout
- Specificity = TN / (TN + FP). Of those without disease, how many test negative. A highly SPecific test, when Positive, rules disease in — SpPin
Sensitivity and specificity are properties of the test and do not change with prevalence.
- Positive predictive value = TP / (TP + FP). Of those testing positive, how many have disease
- Negative predictive value = TN / (TN + FN)
Predictive values depend entirely on prevalence. The same test used in a screening population and in a symptomatic clinic gives very different PPVs. This is the reason population screening for rare disease generates so many false positives, and it is examined constantly.
Likelihood ratios avoid the problem: LR+ = sensitivity / (1 − specificity). An LR+ above 10 or an LR− below 0.1 shifts probability meaningfully.
Measures of effect
| Measure | Meaning | Used in |
|---|---|---|
| Relative risk | Risk in exposed ÷ risk in unexposed | Cohort, RCT |
| Odds ratio | Odds in cases ÷ odds in controls | Case–control; approximates RR when disease is rare |
| Absolute risk reduction | Control event rate − treatment event rate | RCT |
| Number needed to treat | 1 / ARR | RCT |
Relative risk reduction sounds impressive and hides the absolute benefit. A drop from 2% to 1% is a 50% relative reduction, a 1% absolute reduction, and an NNT of 100. When a stem gives both, the absolute figure is usually the point.
Study designs, strongest to weakest
Systematic review of RCTs → RCT → cohort → case–control → cross-sectional → case series → expert opinion.
- Cohort: exposure → outcome, forward in time. Gives incidence and relative risk. Poor for rare outcomes
- Case–control: outcome → exposure, backwards. Efficient for rare diseases; gives odds ratios; vulnerable to recall bias
- Cross-sectional: a snapshot. Gives prevalence; cannot establish that cause preceded effect
Bias and confounding
- Selection bias — who ends up in the study
- Recall bias — cases remember exposures better; the characteristic weakness of case–control studies
- Observer/measurement bias — addressed by blinding
- Lead time bias — screening moves diagnosis earlier without extending life
- Length time bias — screening preferentially detects slow-growing disease
- Confounding — a third factor associated with both exposure and outcome. Controlled by randomisation, restriction, matching, stratification or multivariable adjustment
- Intention to treat preserves randomisation and protects against attrition bias; per-protocol analysis exaggerates benefit
Significance and precision
A p-value is the probability of results this extreme if the null hypothesis were true. It says nothing about effect size or clinical importance.
Confidence intervals are more useful. A 95% CI crossing 1 for a ratio measure — or 0 for a difference — is not statistically significant. A wide interval means an imprecise study, usually an underpowered one.
Type I error (α) is a false positive; type II error (β) a false negative. Power = 1 − β, and is increased mainly by increasing the sample size.