The HAM-D (also written HDRS) grades the severity of a depressive episode in patients who already have a diagnosis of depression. A trained clinician rates 17 items from a clinical interview covering the past week: depressed mood, guilt, suicidal thoughts, three forms of insomnia, work and interests, psychomotor retardation, agitation, psychic and somatic anxiety, gastrointestinal and general somatic symptoms, genital symptoms, hypochondriasis, weight loss and insight.
It is an observer-rated severity measure, not a self-report screener. The rating reflects the clinician's judgement across the whole interview, supplemented where useful by information from relatives or nursing staff. Because roughly half of the possible points sit on somatic and sleep items, the total is best read as a broad index of episode severity rather than a pure measure of mood.
02 - Origin & purpose
Where it comes from.
Max Hamilton published the scale in 1960 while working at the University of Leeds, intending it as a way of quantifying the results of treatment in patients already diagnosed with depressive illness - explicitly not as a diagnostic instrument. He revised the guidance in 1967. The 17-item version scores the severity total; the additional items found on longer versions (diurnal variation, depersonalisation, paranoid symptoms, obsessional symptoms) describe the episode but are not added to the total.
The HAM-D became the reference outcome measure of the antidepressant era: most regulatory trials of antidepressants from the 1960s onwards reported HAM-D change as their primary endpoint, and newer scales such as the MADRS were validated against it. The original scale is in the public domain. Structured interview guides that standardise its administration, such as the SIGH-D and the GRID-HAMD, carry separate copyright and are not reproduced here.
03 - Scoring & cutoffs
How scoring works.
Nine items are rated 0-4 and eight items 0-2, giving a total of 0 to 52. The traditional severity bands shown in the table (0-7 none, 8-13 mild, 14-18 moderate, 19-22 severe, 23+ very severe) are widely used conventions. An empirical study by Zimmerman and colleagues (2013), which anchored HAM-D scores against a structured severity interview in 627 outpatients with major depression, supported a threshold of 7 or below for no depression, 8-16 for mild, 17-23 for moderate and 24 or above for severe depression - slightly broader bands than the convention. In trials and clinical practice, response is usually defined as a reduction of at least 50 percent from baseline, and remission as a total score of 7 or below.
Score
Severity
Interpretation
0–7
No depression
No depression. Symptoms within the normal range.
8–13
Mild
Mild depression.
14–18
Moderate
Moderate depression. A treatment plan should be considered.
19–22
Severe
Severe depression. Active treatment is indicated.
23–52
Very severe
Very severe depression. Active treatment and close risk review are indicated.
04 - Validation evidence
How well it performs.
The HAM-D is one of the most extensively studied instruments in psychiatry. A meta-analysis by Trajković and colleagues covering 409 reliability studies published between 1960 and 2008 found good overall internal consistency and excellent inter-rater reliability when raters are trained, though reliability at the level of individual items is weaker - the insight item in particular performs poorly. Test-retest reliability falls as the interval between ratings lengthens, which is expected for a scale designed to detect change in an evolving illness.
α ≈ 0.79
Internal consistency
Pooled Cronbach's alpha across studies, 1960-2008 (95% CI 0.77-0.81) (Trajković et al., 2011)
ICC ≈ 0.94
Inter-rater reliability
Pooled intraclass correlation for trained raters (95% CI 0.91-0.95) (Trajković et al., 2011)
0.65-0.98
Test-retest reliability
Range across studies; decreases as the retest interval lengthens (Trajković et al., 2011)
17 / 24
Empirical severity cutoffs
ROC-derived thresholds for moderate and severe depression in 627 outpatients (Zimmerman et al., 2013)
05 - How it compares
How it compares to the alternatives.
Instrument
Items
Time
When to reach for it
HAM-D (this page)
17
20-30 min clinician interview
You need the historical reference standard for grading severity and tracking treatment response in diagnosed depression.
You are assessing older adults and want to minimise somatic items that overlap with physical illness.
06 - When to use it
Right tool, wrong tool.
Use it to grade severity and track change over time in patients with an established diagnosis, for example before and after starting treatment. It is clinician-rated from an interview, so it is not suitable as a patient send-out. For self-report screening, the PHQ-9 is a better fit.
Reach for it when
-Grading the severity of a diagnosed depressive episode from a clinical interview.
-Tracking treatment response in settings that follow the trial literature - the response (≥50 percent reduction) and remission (≤7) conventions come from HAM-D research.
-Comparing outcomes against the older antidepressant evidence base, which reports HAM-D endpoints.
-Research or audit where a clinician-rated measure is required and raters can be trained.
Reach for something else when
-Screening undiagnosed patients - it assumes a diagnosis has already been made; use the PHQ-9 or PHQ-2 instead.
-Patient send-outs or waiting-room use - it cannot be self-completed.
-Populations where somatic items mislead, such as significant physical illness - the somatic weighting inflates scores; consider the MADRS or GDS-15.
-Quick reviews without time for a structured interview - a properly rated HAM-D takes 20-30 minutes.
07 - Confidence & precision
Reading the score with care.
Change on the HAM-D should be read against what clinicians can actually detect. Leucht and colleagues linked HAM-D-17 change to clinician global impressions and found that a change of up to 3 points corresponds to "no change" on the CGI - differences that small sit within measurement noise. A change of around 7 points corresponds to "minimally improved", and around 14 points to "much improved". Estimates of the minimal important difference for the HAM-D-17 range from 3 to 8 points, with the most credible values between 3 and 5; NICE has used a 3-point drug-placebo difference as a criterion for clinical significance. In practice: treat movements of 1-3 points as unremarkable, and look for shifts of several points, or crossing the remission threshold, before concluding that a patient has changed.
08 - Limitations
What it cannot tell you.
It reflects a 1960 conception of depression: heavy weighting of insomnia and somatic symptoms, no items for atypical features such as hypersomnia or weight gain, and no coverage of cognitive symptoms as understood today.
Bagby and colleagues' systematic review (2004) concluded the scale is multidimensional, that several items have poor reliability at item level, and that the total conceals which symptoms changed - they argued the "gold standard" had become a lead weight.
Reliability depends on rater training; unstructured administration degrades inter-rater agreement, which is why structured guides (SIGH-D, GRID-HAMD) exist.
Somatic weighting inflates scores in patients with physical illness and in older adults.
It assumes an established diagnosis; it has no validated role as a screening instrument.
Suicide is covered by a single item - the HAM-D is not a risk assessment tool.
See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.
We value your privacy
We use cookies to analyse site usage and improve your experience. Analytics and embedded media (e.g. YouTube) only load if you accept. Read our cookie policy.