Skip to main content

    MDI vs Hamilton: I've seen thousands of numbers. Here's my humble take.

    A patient scored just below 50 on the MDI. The next day I rated them 8 on the Hamilton. After thousands of these questionnaires, here is what I think a screening score can and cannot tell you.

    MDI vs Hamilton: I've seen thousands of numbers. Here's my humble take.
    Nicolas HansenChief physician, Psykiatrien Øst, Region Zealand
    Published 6 August 2026rating scales

    A referral that changed how I view a psychometric questionnaire

    A general practitioner sends a Major Depression Inventory (MDI) home with a patient. The patient completes it at the kitchen table and posts it back. The score comes in just below 50, high enough that the patient is told to present to the psychiatric ER. I saw them the following day.

    I remember one specific case vividly. I rated the same patient at 8 on the Hamilton Depression Rating Scale.

    They were genuinely burdened. Life had been hard on them and they needed help. But when I worked through the clinical picture systematically, there was nothing depressive in it. No melancholic core, no diurnal variation, no anhedonia in the technical sense. The quality of symptoms did not in my clinical evaluation reach malign proportions. A person under strain, correctly identified as suffering and incorrectly identified as depressed.

    I spent roughly five to six years in an acute psychiatric admissions unit, three of them as chief of department, with somewhere between ten and sixteen new admissions a day, roughly two hundred workdays a year. I have seen thousands of MDIs. That mismatch, an alarming self-report score alongside a modest clinician rating, came through the doors far too often for me to file it as an anomaly.

    This is not an argument that the Major Depression Inventory is a bad instrument. It is an argument to not rely solely on numbers, but always to validate them with our clinical knowledge.

    What the Major Depression Inventory is, and what it does well

    The Major Depression Inventory is a ten-item self-report questionnaire developed in Denmark by Per Bech and colleagues, scored 0 to 50, that can be read two ways: as a diagnostic algorithm that leads onto ICD-10 and DSM-IV criteria for depressive episode, or as a straightforward severity score. It is free, brief, and available in Danish, which is exactly why it became the default depression questionnaire in Danish general practice.

    Let me be fair to it, because it does several things genuinely well.

    It is enormously time-saving. You hand it to the patient in the waiting room, or send it home with them, and it costs the clinic essentially nothing in staff time. It involves the patient in their own assessment, and patients feel heard. That is not a small thing, and it is not merely cosmetic. A patient who has put their own experience into structured words often arrives at the consultation better prepared to describe it.

    And it casts a wide net. That is by design. In the original validation work against the Schedule for Clinical Assessment in Neuropsychiatry, Bech and colleagues reported sensitivity of 0.86 to 0.92 for the diagnostic algorithms, with specificity of 0.82 to 0.86, and identified 26 as the optimal cutoff on the total score. If someone in your waiting room has moderate to severe depression, an MDI will very rarely miss them.

    Where it falls down: the specificity problem

    The trouble is the other side of the coin. High sensitivity tends to come at the cost of specificity, and a net that wide catches things that are not fish.

    Those original validation figures came from a sample enriched for depression, assessed under research conditions. Move the instrument into an unselected clinical population and the numbers change considerably. In a study of 258 outpatients referred to mental health care, Cuijpers, Dekker, Noteboom, Smits and Peen found the MDI achieved a sensitivity of 0.66 and a specificity of 0.63 against psychiatrist diagnosis, with agreement of only kappa = 0.26. Their conclusion was blunt: the MDI performs adequately as a severity measure, but its diagnostic algorithm should not be relied upon in outpatient populations.

    That is the statistical shape of what I was seeing clinically. Roughly a third of the patients the instrument flagged were not, on physician assessment, depressed.

    There is an inherent limitation to self-report that no amount of validation removes. If you are severely depressed, there may be symptoms you over-report and others you under-report, and there can be many reasons for this, none of them dishonest. Distress, exhaustion, grief, insomnia, financial fear and a marriage coming apart all read as depression on paper. The instrument has no way to ask the follow-up question that separates them. A clinician does.

    What arrived on my desk was a number without a context. And a number without a context, if it is high enough, triggers a pathway: an acute referral, a bed, an assessment, sometimes a prescription. There is a real cost to the service, and more importantly to a patient who has just been told that their score means something it does not.

    The self-rating trap

    There is a subtler problem, and it matters increasingly as clinics adopt digital monitoring.

    The intuition behind frequent symptom tracking is appealing: more data points, earlier detection, better care. It was tested properly here in Denmark. In the MONARCA I trial, Faurholt-Jepsen and colleagues at the Copenhagen Affective Disorder Research Centre randomised 78 patients with bipolar disorder to six months of daily smartphone self-monitoring with a clinical feedback loop, or to a smartphone used normally. Symptoms were tracked with the HAM-D17 and the Young Mania Rating Scale, both clinician-rated instruments, used to evaluate a self-rating intervention.

    Daily self-monitoring did not reduce depressive or manic symptoms. If anything it went the other way. Patients using the MONARCA system had more sustained depressive symptoms than patients using a smartphone for ordinary communication, though they had fewer manic symptoms during the trial. The authors' own conclusion was that electronic self-monitoring, "although intuitive and appealing, needs critical consideration and further clarification before it is implemented as a clinical tool."

    Handing someone a mirror every morning is not a neutral act. If you are depressed and you rate yourself daily, the daily attention to symptoms can make the burden heavier. I keep this study close whenever someone proposes that the answer to a measurement problem is more measurement. It is also why I think between-visit patient monitoring has to be designed with a clinical rationale for every question asked, and a defined stopping point, rather than simply switched on because the data is available.

    Why I advocate for the Hamilton

    The Hamilton Depression Rating Scale is clinician-rated. That is its entire strength, and everything else follows from it.

    It is a structured interview, which means a clinician is present, continuously validating what the patient reports by probing, contextualising, and weighing one answer against the last. When a patient says they are not sleeping, I can establish whether that is initial insomnia, early morning waking, or a newborn in the next room. The questionnaire cannot.

    Within the Hamilton family, the evidence favours brevity. The six-item melancholia subscale (HAM-D6), covering depressed mood, work and interests, general somatic symptoms, psychic anxiety, guilt, and psychomotor retardation, is the part with the strongest psychometric standing. It is the only version that fits a unidimensional Rasch model, meaning its total score behaves as a genuine measure of a single underlying severity continuum rather than a sum of loosely related complaints. It is also at least as sensitive to treatment effect as the full 17-item version, and in antidepressant trials often more so. The classic HAM-D17 has some years on it and is criticised for its somatic and anxiety items, but it covers the ground and remains the reference measure in much of the literature.

    The Hamilton is not diagnostic, and nobody should pretend otherwise. It is a tell tale guide to depressive depth and a genuinely good instrument for monitoring change over time. Its weakness is the obvious one: put two clinicians in front of the same patient and you may get two different ratings. Inter-rater variability is real, and it is the honest counterargument to everything I have written above.

    But I would rather have a rating that a trained clinician can defend and interrogate than an unvalidated number that arrived in the post.

    MDI vs Hamilton: how I actually choose

    MDIHAM-D6 / HAM-D17
    Who completes itPatientClinician, via structured interview
    Time5 minutes, unsupervised15 to 30 minutes of clinical time
    Range0 to 50, cutoff 26 for depression0 to 50 (HAM-D17), 0 to 22 (HAM-D6)
    Best atCasting a wide net, giving the patient a voiceEstablishing depressive depth, tracking treatment response
    Weakest atSpecificity in unselected populationsRater consistency, cost in clinician time
    Use it toDecide who to talk toDecide what is actually going on

    From the Aisel editors: neighbouring instruments

    Nicolas's comparison above is between the two scales in his toolbox. For readers weighing up alternatives, these are the instruments that most often come up in the same decision, and are recommended by DPMG.

    • PHQ-9: the international self-report workhorse, better validated across languages than the MDI and the usual default outside Denmark. Its two-item short form, the PHQ-2, works as a first-pass filter.
    • HADS: built for general medical settings, omitting somatic items that would otherwise confound the score in physically unwell patients.
    • MADRS: the usual clinician-rated alternative to the Hamilton, and often preferred for tracking treatment effect.
    • EPDS: the standard choice in pregnancy and the postnatal year, where the somatic items in general depression scales mislead badly.
    • CGI-S: when the diagnosis is settled and the question is simply global severity, one rating in under a minute.
    • GAF: for function rather than symptoms, with the caveats about reliability set out on that page.

    The full clinical scale library has scoring, cutoffs and validation evidence for each.

    The practical takeaway

    This is not an argument for abandoning self-report scales. It is an argument for knowing what each tool is for.

    A self-report instrument like the Major Depression Inventory is a net: cheap, wide, sensitive. Use it to make sure nobody slips through, and to give patients a structured way to tell you what they are experiencing.

    A clinician-rated instrument like the Hamilton is a scalpel. Use it to establish what the net actually caught, and to follow the patient over time.

    The mistake is treating a screening score as a conclusion. A high MDI is a reason to talk to the patient. It is never a substitute for doing so.

    References

    1. Bech P, Rasmussen NA, Olsen LR, Noerholm V, Abildgaard W (2001). The sensitivity and specificity of the Major Depression Inventory, using the Present State Examination as the index of diagnostic validity. Journal of Affective Disorders, 66(2/3), 159-164. https://doi.org/10.1016/S0165-0327(00)00309-8
    2. Cuijpers P, Dekker J, Noteboom A, Smits N, Peen J (2007). Sensitivity and specificity of the Major Depression Inventory in outpatients. BMC Psychiatry, 7, 39. https://doi.org/10.1186/1471-244X-7-39
    3. Faurholt-Jepsen M, Frost M, Ritz C, et al. (2015). Daily electronic self-monitoring in bipolar disorder using smartphones, the MONARCA I trial. Psychological Medicine, 45(13), 2691-2704. https://doi.org/10.1017/S0033291715000410
    4. Østergaard SD, Bech P, Miskowiak KW (2017). Rasch analysis of the HAM-D6. PLOS ONE, 12(1), e0170000. https://doi.org/10.1371/journal.pone.0170000
    5. Hamilton M (1960). A rating scale for depression. Journal of Neurology, Neurosurgery, and Psychiatry, 23, 56-62. https://doi.org/10.1136/jnnp.23.1.56
    Nicolas Hansen, Chief physician, Psykiatrien Øst, Region Zealand · Clinical adviser to Aisel

    About the author

    Nicolas Hansen

    Chief physician, Psykiatrien Øst, Region Zealand · Clinical adviser to Aisel

    Nicolas Hansen is a chief physician at Psykiatrien Øst, Region Zealand, a role he has held since 2021, following a period as lead consultant. He qualified as a physician in 2009 and has been a specialist in psychiatry since 2018, including three years in acute psychiatric admissions.

    Read the full profile

    Ready to transform your clinical workflow?

    See how Aisel can automate intake, documentation, and EHR sync for your practice.

    Schedule a Demo