The HAM-D measures the severity of depressive symptoms in people who already carry a depression diagnosis. It is clinician-rated, not self-report: a trained rater interviews the patient about the previous week and scores each item from observation and report combined. Nine items are rated 0–4 and eight 0–2, giving a total of 0–52 on the standard 17-item form.
Its coverage is deliberately clinical rather than diagnostic. Mood, guilt, suicidal thinking, work and activities, retardation and agitation carry the heaviest weight, alongside anxiety and a substantial block of somatic and sleep items. The 21-item version adds four further items intended to subtype the depression rather than grade its severity, which is why the 17-item form remains the standard severity measure.
02 - Origin & purpose
Where it comes from.
Max Hamilton published the scale in 1960 in the Journal of Neurology, Neurosurgery and Psychiatry, designed to quantify severity in patients already diagnosed with depressive illness so that treatment effects could be measured. He revised and factor-analysed it in 1967. It was never intended as a screening or case-finding instrument, and it still should not be used as one.
The HAM-D became the primary efficacy endpoint in antidepressant trials and has held that position for six decades, which is the main argument for continuing to use it: comparability with a very large historical evidence base. Circulating versions of the form differ from Hamilton's original in their anchors, so record which version you used. Structured interview guides - the SIGH-D (Williams, 1988) and the GRID-HAMD (Williams et al., 2008) - were developed to reduce the variability that free-form administration introduces.
03 - Scoring & cutoffs
How scoring works.
The rater sums all 17 items for a total of 0–52. Remission is conventionally defined as a total of 7 or below, and treatment response as a reduction of 50% or more from baseline; both conventions are long established in the trial literature.
The severity bands need a caveat. The bands shown below are the conventional textbook set, but they were never derived from a single validation study. Zimmerman and colleagues (2013) derived bands empirically against clinician global severity ratings and found different boundaries: 0–7 no depression, 8–16 mild, 17–23 moderate, and 24 or above severe. The practical difference sits in the high teens and low twenties: a total of 20 is "severe" on the conventional scheme and "moderate" on the empirical one. Where the distinction matters, state which scheme you are using.
Score
Severity
Interpretation
0–7
No depression
No depression. Symptoms within the normal range.
8–13
Mild
Mild depression.
14–18
Moderate
Moderate depression. A treatment plan should be considered.
19–22
Severe
Severe depression. Active treatment is indicated.
23–52
Very severe
Very severe depression. Active treatment and close risk review are indicated.
04 - Validation evidence
How well it performs.
A reliability meta-analysis by Trajković and colleagues (2011) pooled 409 reports published between 1960 and 2008. Internal consistency was moderate rather than high (pooled Cronbach's alpha 0.79, 95% CI 0.77–0.81), while inter-rater agreement on the total score was excellent (pooled ICC 0.94, 95% CI 0.91–0.95). Test-retest coefficients ranged from 0.65 to 0.98 and fell as the interval lengthened.
Bagby and colleagues' 2004 review is the necessary counterweight. Agreement on the total score is reliably good; agreement at item level is not. Several individual items - loss of insight, hypochondriasis, somatic anxiety, general somatic symptoms and agitation - show poor reliability, and the scale's factor structure does not replicate: across 17 samples, analyses extracted anywhere between two and eight factors. Convergent validity with the MADRS is strong (r = 0.68–0.88).
0.79
POOLED CRONBACH'S α
0.94
INTER-RATER ICC
≤7
REMISSION THRESHOLD
3–5 pts
MINIMAL IMPORTANT DIFFERENCE
05 - How it compares
How it compares to the alternatives.
Instrument
Items
Time
When to reach for it
HAM-D
17
~20 min
Clinician-rated severity where comparability with the antidepressant trial literature matters. Public domain.
Medical and hospital settings: deliberately excludes somatic items and reports anxiety alongside depression.
06 - When to use it
Right tool, wrong tool.
Use it to grade severity and track change over time in patients with an established diagnosis, for example before and after starting treatment. It is clinician-rated from an interview, so it is not suitable as a patient send-out. For self-report screening, the PHQ-9 is a better fit.
Reach for it when
-Grading severity in a patient already diagnosed with a depressive disorder.
-Running or reading a treatment trial where HAM-D comparability is expected.
-You need clinician observation rather than self-report, for example where insight is limited or self-report seems unreliable.
-Documenting response (≥50% reduction) or remission (≤7) for a specialist care record.
Reach for something else when
-You are screening an undiagnosed patient: the HAM-D was never built for case-finding; use the PHQ-9.
-The patient is medically ill and somatic and sleep items will inflate the total; use the MADRS or HADS.
-No trained rater is available; unstructured administration degrades reliability. Use a self-report measure instead.
-You are assessing a child or adolescent; the HAM-D is not validated for that population.
-You need to detect atypical features such as hypersomnia and hyperphagia, which the 17-item form barely covers.
07 - Confidence & precision
Reading the score with care.
No standard error of measurement for the 17-item HAM-D is established in the published literature, so we do not quote one. What is documented is the minimal important difference: Hengartner and Plöderl (2022) reviewed anchor- and distribution-based estimates and put the range at 3 to 8 points, with 3 to 5 points the most defensible. NICE has used a 3-point drug-placebo difference as its threshold of clinical relevance.
Two practical consequences. First, changes of one or two points between visits should not be read as improvement or deterioration. Second, because item-level agreement between raters is weak even where total-score agreement is strong, serial ratings are most interpretable when the same rater, using the same structured guide, performs them.
08 - Limitations
What it cannot tell you.
The HAM-D is a severity instrument for diagnosed depression, not a diagnostic or screening tool. Its content is unbalanced: three separate insomnia items sit alongside thin coverage of concentration difficulty and worthlessness, and somatic weighting inflates scores in physically unwell patients. It is not unidimensional, so a single total conflates distinct symptom domains, and the same total can be reached by very different clinical pictures. Atypical presentations are poorly captured. Administration varies between raters unless a structured guide is used, and multiple non-identical versions of the form circulate. Bagby and colleagues put the case bluntly in 2004, asking whether the gold standard had become a lead weight, a fair summary of why MADRS and self-report measures have displaced it in much routine practice.
Get our case study pack by email, with practical examples from clinical teams.
We will only use your details to send relevant Aisel resources. Unsubscribe at any time.
We value your privacy
We use cookies to analyse site usage and improve your experience. Analytics and embedded media (e.g. YouTube) only load if you accept. Read our cookie policy.