In the only blinded randomised trial of its kind, outpatients with major depression whose treatment was guided by rating scales remitted at nearly three times the rate of those receiving standard care - 74% against 29% 1. Unstructured clinical judgement misses deterioration, delays dose adjustments, and lets non-response drift on for months. Self-report scales fix this cheaply: patients disclose more on a form than face to face, and the score arrives before the consultation rather than consuming it.
There is also a structural reason this is no longer optional. Psychiatry is moving towards outcome-based care - commissioners, insurers and referrers increasingly want clinics to demonstrate results, not just activity 2. You cannot run outcome-based care on clinical impressions; it needs a validated number, captured routinely, at near-zero cost per administration. That specification describes a free self-report scale.
The obstacle is rarely conviction; it's cost and habit. Several of the best-known instruments are not free - the BDI-II is licensed by Pearson, the HADS too - and clinicians assume licensed must mean better. The evidence says otherwise. This shortlist covers what you can use today without a fee, and when to reach for each.
The shortlist
| Scale | Items | Best for | Cutoff |
|---|---|---|---|
| PHQ-9 | 9 | Default screening and monitoring | ≥10 |
| PHQ-2 | 2 | Universal screening gate (≥2 → PHQ-9) | ≥2 |
| MDI | 10-12 | Diagnosis-aligned scoring (ICD-10/DSM) | Severity bands |
| GDS-15 | 15 | Older adults | ≥5 |
| EPDS | 10 | Perinatal | ≥11 (≥13 for higher specificity) |
| CES-D | 20 | Research continuity | ≥16 |
1. PHQ-9 - the default
Nine items mapped onto the DSM criteria for major depression, scored 0-27, free to reproduce. The DEPRESSD individual-participant-data meta-analysis - 29 studies and 6,725 participants assessed with semi-structured interviews - put sensitivity at 0.88 and specificity at 0.85 at the standard cutoff of ≥10 3. It also has well-characterised sensitivity to change, which makes it the sensible default for both intake and session-to-session monitoring. Item 9 (suicidal ideation) needs a follow-up pathway; where that cannot be guaranteed, the PHQ-8 drops the item and keeps the same cutoff. Full brief →
2. PHQ-2 - for universal screening, not specialist intake
The evidence for the two-step approach is real, but it belongs in a specific place. In the DEPRESSD meta-analysis, a PHQ-2 score of ≥2 followed by a full PHQ-9 matched the sensitivity of the PHQ-9 alone with higher specificity 4 - a good trade in primary care, oncology or antenatal programmes, where most patients screen negative and volume matters. Note the gate is ≥2, not the commonly quoted ≥3, which costs sensitivity. In a psychiatric clinic, where pre-test probability is high and the questionnaire is self-completed digitally, skip the gate and use the PHQ-9 directly. Full brief →
3. MDI - when you want diagnosis, not just severity
The Major Depression Inventory (Bech, WHO lineage) is the only free instrument here that yields both a severity score and a diagnostic algorithm aligned to ICD-10 and DSM criteria. It is the standard in Danish practice and underused elsewhere. Worth knowing: it produces more false positives than the Hamilton in some settings - we've written about that here. Full brief →
4. GDS-15 - older adults
Fifteen yes/no items designed for older adults, deliberately avoiding somatic symptoms that overlap with ageing and physical illness; five or more suggests depression. The binary format also suits patients who struggle with graded response options. Full brief →
5. EPDS - perinatal
The Edinburgh Postnatal Depression Scale is the perinatal standard, free to reproduce with citation. General depression scales miss the mark here because sleep disturbance and fatigue are near-universal postpartum; the EPDS avoids them. On cutoffs, the individual-participant-data meta-analysis is worth internalising: ≥11 maximises combined accuracy (sensitivity 0.81, specificity 0.88), while the traditional ≥13 is far more specific but misses a third of cases (sensitivity 0.66, specificity 0.95) 5. Choose by what a false negative costs in your pathway. Full brief →
6. CES-D - the research workhorse
Radloff's 1977 scale, public domain, with decades of epidemiological data behind the ≥16 cutoff. Choose it when comparability with existing research cohorts matters; for routine clinical work the PHQ-9 is briefer and more actionable. Full brief →
What to avoid paying for
This is a sourced claim, not a preference. The HADS-D, at its accuracy-maximising cutoff of ≥7, reaches sensitivity 0.82 and specificity 0.78 6 - lower than the PHQ-9 on both counts, from the same research programme using the same methodology. The BDI-II performs respectably: it correlates strongly with the PHQ-9 (r = 0.77) 7 and shows similar responsiveness to change during treatment 8 - but "similar for a fee" is not an argument, and comparison authors note the PHQ-9's brevity and direct mapping to diagnostic criteria tilt the advantage the other way. The HAM-D and MADRS are excellent instruments but clinician-rated - a different job entirely.
The bottom line
Use the PHQ-9 as the default, and let it do double duty: the same score that screens at intake becomes the outcome data your clinic will increasingly be asked to show. Add the PHQ-2 gate only in universal-screening programmes. Swap in the GDS-15 for older adults, the EPDS perinatally, and the MDI when you want diagnosis-aligned output. Every scale above is free, validated, and available with full scoring guidance in the Aisel scale library.
References
- Guo T, Xiang YT, Xiao L, et al. Measurement-based care versus standard care for major depression: a randomized controlled trial with blind raters. American Journal of Psychiatry. 2015;172(10):1004-1013. https://doi.org/10.1176/appi.ajp.2015.14050652 ↩
- Fortney JC, Unützer J, Wrenn G, et al. A tipping point for measurement-based care. Psychiatric Services. 2017;68(2):179-188. https://doi.org/10.1176/appi.ps.201500439 ↩
- Levis B, Benedetti A, Thombs BD; DEPRESSD Collaboration. Accuracy of the Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis. BMJ. 2019;365:l1476. https://doi.org/10.1136/bmj.l1476 ↩
- Levis B, Sun Y, He C, et al. Accuracy of the PHQ-2 alone and in combination with the PHQ-9 for screening to detect major depression. JAMA. 2020;323(22):2290-2300. https://doi.org/10.1001/jama.2020.6504 ↩
- Levis B, Negeri Z, Sun Y, et al. Accuracy of the Edinburgh Postnatal Depression Scale (EPDS) for screening to detect major depression among pregnant and postpartum women: individual participant data meta-analysis. BMJ. 2020;371:m4022. https://doi.org/10.1136/bmj.m4022 ↩
- Wu Y, Levis B, Sun Y, et al. Accuracy of the Hospital Anxiety and Depression Scale Depression subscale (HADS-D) to screen for major depression: individual participant data meta-analysis. BMJ. 2021;373:n972. https://doi.org/10.1136/bmj.n972 ↩
- Kung S, Alarcon RD, Williams MD, et al. Comparing the Beck Depression Inventory-II (BDI-II) and Patient Health Questionnaire (PHQ-9) depression measures in an integrated mood disorders practice. Journal of Affective Disorders. 2013;145(3):341-343. https://doi.org/10.1016/j.jad.2012.08.017 ↩
- Titov N, Dear BF, McMillan D, et al. Psychometric comparison of the PHQ-9 and BDI-II for measuring response during treatment of depression. Cognitive Behaviour Therapy. 2011;40(2):126-136. https://doi.org/10.1080/16506073.2010.550059 ↩
Photo: Nguyen Dang Hoang Nhu on Unsplash.
