The PHQ-2 is an ultra-brief depression screen consisting of the first two items of the PHQ-9: little interest or pleasure in doing things, and feeling down, depressed or hopeless. These are the two cardinal symptoms of major depression - anhedonia and low mood - one of which is required for the diagnosis in both DSM and ICD. Each item is rated 0-3 for frequency over the last two weeks, giving a total of 0-6.
It answers exactly one question: is depression likely enough in this person to justify a fuller assessment? It grades nothing, tracks nothing and diagnoses nothing. A positive screen should be followed by the full PHQ-9 or a clinical interview; a negative screen in a low-risk population makes depression unlikely.
02 - Origin & purpose
Where it comes from.
Kroenke, Spitzer and Williams published the PHQ-2 in 2003, validating it in 580 patients whose structured mental-health interviews had been obtained in the original PHQ primary-care and obstetrics-gynaecology studies 1. Receiver-operating-characteristic analysis identified 3 as the optimal cutpoint, at which sensitivity was 0.83 and specificity 0.92 for major depression.
The purpose was pragmatic: universal screening fails when the instrument is too long for routine use. Two questions can be asked verbally in under a minute or embedded in any intake form, making the PHQ-2 the standard first stage of two-step screening programmes in primary care, oncology, antenatal care and general medicine. Like the whole PHQ family it is free to reproduce without permission.
03 - Scoring & cutoffs
How scoring works.
Both items are summed to a total of 0-6. The traditional cutoff is ≥3. The best current evidence complicates that slightly: in the DEPRESSD individual-participant-data meta-analysis, ≥3 had sensitivity 0.72 - it misses more than a quarter of major depression - while ≥2 had sensitivity 0.91 at the cost of specificity (0.67) 2. In a two-step programme where positives go on to the PHQ-9, the lower gate of ≥2 is the better choice: the second step absorbs the false positives, and the combination matched the sensitivity of giving everyone a PHQ-9 with higher overall specificity. Use ≥3 only where a positive screen triggers a costly response directly.
Score
Severity
Interpretation
0–2
Negative screen
Below the cutoff. Depression is unlikely on this screen.
3–6
Positive screen
At or above the cutoff of 3. Follow up with the PHQ-9 or a clinical interview.
04 - Validation evidence
How well it performs.
The original validation put sensitivity at 0.83 and specificity at 0.92 at ≥3 against mental-health-professional interview 1. The DEPRESSD collaboration's individual-participant-data meta-analysis (100 studies, over 44,000 participants) gave more conservative estimates against semi-structured diagnostic interviews: sensitivity 0.72 and specificity 0.85 at ≥3, versus 0.91 and 0.67 at ≥2 - and showed that PHQ-2 (≥2) followed by PHQ-9 (≥10) matched the sensitivity of PHQ-9 alone with higher specificity 2. A separate diagnostic meta-analysis reached the same conclusion about the cutoff trade-off 3. Internal consistency is acceptable for a two-item scale (Cronbach's alpha about 0.75-0.79 across studies) and one-week test-retest reliability around 0.76 has been reported 4.
0.91
SENSITIVITY (≥2)
0.67
SPECIFICITY (≥2)
0.72
SENSITIVITY (≥3)
≈0.75-0.79
CRONBACH'S α
05 - How it compares
How it compares to the alternatives.
Instrument
Items
Time
When to reach for it
PHQ-2
2
<1 min
Ultra-brief two-item gate; screens only, no severity grading.
Older adults, where somatic confounds and graded options are a problem.
06 - When to use it
Right tool, wrong tool.
Reach for it when
-Universal screening in high-volume, low-prevalence settings (primary care, antenatal clinics, oncology, general medicine).
-Verbal screening in consultations or on the phone.
-The first stage of a two-step programme with the PHQ-9 as the second stage.
-Intake forms where every item counts.
Reach for something else when
-Specialist psychiatric intake, where pre-test probability is high - go straight to the PHQ-9.
-Severity grading or monitoring treatment response.
-Any setting where a positive screen will not be followed up - a screen without a second step is theatre.
-Suicide-risk assessment - neither item asks about it.
07 - Confidence & precision
Reading the score with care.
With a 0-6 range, the PHQ-2 has no meaningful severity gradations, no established standard error of measurement and no minimal clinically important difference - nor should it: it is a binary gate, not a measure. Precision is a property of the cutoff chosen. At ≥2, about one depressed patient in eleven screens negative; at ≥3, more than one in four does. Interpret a borderline score by asking the two questions again in conversation, not by arithmetic.
08 - Limitations
What it cannot tell you.
Sensitivity at the traditional ≥3 cutoff is materially lower than commonly assumed (0.72 in the best meta-analytic evidence), so programmes using ≥3 as a hard gate will miss depressed patients. Specificity at ≥2 is poor in isolation, so the PHQ-2 only works as part of a two-step pathway. It contains no suicidality item, cannot grade severity or track change, and - like all the PHQ instruments - is a self-report screen whose positive results require diagnostic confirmation. Accuracy estimates come mostly from general-medical rather than specialist psychiatric populations.
See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.
We value your privacy
We use cookies to analyse site usage and improve your experience. Analytics and embedded media (e.g. YouTube) only load if you accept. Read our cookie policy.