Skip to main content
    Depression · 2 items · 0–6 · Kroenke K, Spitzer RL, Williams JBW (2003). Med Care.

    Patient Health Questionnaire-2: Scoring, Cutoffs & Interpretation

    Two-item ultra-brief screen for depression - the first two questions of the PHQ-9.

    PHQ-20 / 2

    Over the last two weeks, how often have you been bothered by…

    Scored locally - nothing leaves this page

    01Little interest or pleasure in doing things.
    02Feeling down, depressed, or hopeless.
    0 of 20 / 6

    02 - The clinician's brief

    Last reviewed: · Reviewed by Lotte Kjær Svalberg

    01 - What it measures

    What this scale measures.

    The PHQ-2 is an ultra-brief depression screen consisting of the first two items of the PHQ-9: little interest or pleasure in doing things, and feeling down, depressed or hopeless. These are the two cardinal symptoms of major depression - anhedonia and low mood - one of which is required for the diagnosis in both DSM and ICD. Each item is rated 0-3 for frequency over the last two weeks, giving a total of 0-6.

    It answers exactly one question: is depression likely enough in this person to justify a fuller assessment? It grades nothing, tracks nothing and diagnoses nothing. A positive screen should be followed by the full PHQ-9 or a clinical interview; a negative screen in a low-risk population makes depression unlikely.

    02 - Origin & purpose

    Where it comes from.

    Kroenke, Spitzer and Williams published the PHQ-2 in 2003, validating it in 580 patients whose structured mental-health interviews had been obtained in the original PHQ primary-care and obstetrics-gynaecology studies 1. Receiver-operating-characteristic analysis identified 3 as the optimal cutpoint, at which sensitivity was 0.83 and specificity 0.92 for major depression.

    The purpose was pragmatic: universal screening fails when the instrument is too long for routine use. Two questions can be asked verbally in under a minute or embedded in any intake form, making the PHQ-2 the standard first stage of two-step screening programmes in primary care, oncology, antenatal care and general medicine. Like the whole PHQ family it is free to reproduce without permission.

    03 - Scoring & cutoffs

    How scoring works.

    Both items are summed to a total of 0-6. The traditional cutoff is ≥3. The best current evidence complicates that slightly: in the DEPRESSD individual-participant-data meta-analysis, ≥3 had sensitivity 0.72 - it misses more than a quarter of major depression - while ≥2 had sensitivity 0.91 at the cost of specificity (0.67) 2. In a two-step programme where positives go on to the PHQ-9, the lower gate of ≥2 is the better choice: the second step absorbs the false positives, and the combination matched the sensitivity of giving everyone a PHQ-9 with higher overall specificity. Use ≥3 only where a positive screen triggers a costly response directly.

    Score
    Severity
    Interpretation
    0–2
    Negative screen
    Below the cutoff. Depression is unlikely on this screen.
    3–6
    Positive screen
    At or above the cutoff of 3. Follow up with the PHQ-9 or a clinical interview.

    04 - Validation evidence

    How well it performs.

    The original validation put sensitivity at 0.83 and specificity at 0.92 at ≥3 against mental-health-professional interview 1. The DEPRESSD collaboration's individual-participant-data meta-analysis (100 studies, over 44,000 participants) gave more conservative estimates against semi-structured diagnostic interviews: sensitivity 0.72 and specificity 0.85 at ≥3, versus 0.91 and 0.67 at ≥2 - and showed that PHQ-2 (≥2) followed by PHQ-9 (≥10) matched the sensitivity of PHQ-9 alone with higher specificity 2. A separate diagnostic meta-analysis reached the same conclusion about the cutoff trade-off 3. Internal consistency is acceptable for a two-item scale (Cronbach's alpha about 0.75-0.79 across studies) and one-week test-retest reliability around 0.76 has been reported 4.

    0.91
    SENSITIVITY (≥2)
    0.67
    SPECIFICITY (≥2)
    0.72
    SENSITIVITY (≥3)
    ≈0.75-0.79
    CRONBACH'S α

    05 - How it compares

    How it compares to the alternatives.

    Instrument
    Items
    Time
    When to reach for it
    PHQ-2
    2
    <1 min
    Ultra-brief two-item gate; screens only, no severity grading.
    9
    2-3 min
    The step-up instrument: severity grading, change tracking, suicidality item.
    8
    2-3 min
    PHQ-9 without the self-harm item, for settings without a follow-up pathway.
    2
    <1 min
    The anxiety counterpart; together they form the PHQ-4.
    4
    ~1 min
    PHQ-2 + GAD-2 combined when both domains need a gate.
    5
    ~1 min
    Positively phrased wellbeing screen; useful where a deficit-framed screen is unwelcome.
    15
    ~5 min
    Older adults, where somatic confounds and graded options are a problem.

    06 - When to use it

    Right tool, wrong tool.

    Reach for it when

    • -Universal screening in high-volume, low-prevalence settings (primary care, antenatal clinics, oncology, general medicine).
    • -Verbal screening in consultations or on the phone.
    • -The first stage of a two-step programme with the PHQ-9 as the second stage.
    • -Intake forms where every item counts.

    Reach for something else when

    • -Specialist psychiatric intake, where pre-test probability is high - go straight to the PHQ-9.
    • -Severity grading or monitoring treatment response.
    • -Any setting where a positive screen will not be followed up - a screen without a second step is theatre.
    • -Suicide-risk assessment - neither item asks about it.

    07 - Confidence & precision

    Reading the score with care.

    With a 0-6 range, the PHQ-2 has no meaningful severity gradations, no established standard error of measurement and no minimal clinically important difference - nor should it: it is a binary gate, not a measure. Precision is a property of the cutoff chosen. At ≥2, about one depressed patient in eleven screens negative; at ≥3, more than one in four does. Interpret a borderline score by asking the two questions again in conversation, not by arithmetic.

    08 - Limitations

    What it cannot tell you.

    Sensitivity at the traditional ≥3 cutoff is materially lower than commonly assumed (0.72 in the best meta-analytic evidence), so programmes using ≥3 as a hard gate will miss depressed patients. Specificity at ≥2 is poor in isolation, so the PHQ-2 only works as part of a two-step pathway. It contains no suicidality item, cannot grade severity or track change, and - like all the PHQ instruments - is a self-report screen whose positive results require diagnostic confirmation. Accuracy estimates come mostly from general-medical rather than specialist psychiatric populations.

    Common questions from clinicians

    FAQ

    References

    Source literature.

    1. [1]Kroenke K, Spitzer RL, Williams JBW The Patient Health Questionnaire-2: validity of a two-item depression screener. Medical Care. 2003;41(11):1284-1292 (2003)
    2. [2]Levis B, Sun Y, He C, et al. Accuracy of the PHQ-2 alone and in combination with the PHQ-9 for screening to detect major depression: systematic review and meta-analysis. JAMA. 2020;323(22):2290-2300 (2020)
    3. [3]Manea L, Gilbody S, Hewitt C, et al. Identifying depression with the PHQ-2: a diagnostic meta-analysis. Journal of Affective Disorders. 2016;203:382-395 (2016)
    4. [4]Löwe B, Kroenke K, Gräfe K Detecting and monitoring depression with a two-item questionnaire (PHQ-2). Journal of Psychosomatic Research. 2005;58(2):163-171 (2005)
    5. [5]Arroll B, Goodyear-Smith F, Crengle S, et al. Validation of PHQ-2 and PHQ-9 to screen for major depression in the primary care population. Annals of Family Medicine. 2010;8(4):348-353 (2010)

    Scale without compromise

    See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.