The PHQ-4 is a four-item ultra-brief screen for the two commonest mental health presentations in general practice: anxiety and depression. It is a composite rather than a single-construct instrument. The first two items are the GAD-2, the second two are the PHQ-2, and the scale therefore yields two two-item subscale scores alongside a total distress score from 0 to 12. Respondents rate how often each symptom has bothered them over the preceding two weeks, from "not at all" (0) to "nearly every day" (3).
The total score is best read as a general index of anxious and depressive distress rather than as a measure of either condition on its own. A systematic review of 26 studies covering 93,466 participants across 19 countries found the two-factor structure to be the most consistently supported solution, invariant across gender, age, geography, income, education and language. In the German general population the two subscales correlated at r = 0.61 - closely related, but distinguishable. In practice the subscale scores carry the clinical signal: the total tells you how much distress is present, not what kind.
02 - Origin & purpose
Where it comes from.
Kroenke, Spitzer, Williams and Löwe published the PHQ-4 in Psychosomatics in 2009, assembling it from two screeners already in wide use rather than writing new items. Their validation sample comprised 2,149 patients drawn from 15 primary care clinics, aged 18 to 94, with a mean age of 47.2 years. The aim was to compress screening for the two most prevalent disorders into a burden small enough to be administered universally, at every visit if desired, rather than only where a clinician already suspects a problem.
Löwe and colleagues standardised the instrument the following year in a nationally representative German household survey of roughly 5,030 adults, producing percentile norms by age and gender that remain the reference distribution for the scale. A 2022 update of that standardisation found only minor changes. The design has one further practical consequence: because the PHQ-4 nests exactly inside longer instruments - the GAD-2 is the first two items of the GAD-7, the PHQ-2 the first two of the PHQ-9 - a positive screen routes onto the parent instrument without re-asking any question the patient has already answered.
03 - Scoring & cutoffs
How scoring works.
Sum all four items for a total between 0 and 12, then score the two subscales separately. Items 1 and 2 (feeling nervous or on edge; being unable to stop or control worrying) form the anxiety subscale. Items 3 and 4 (little interest or pleasure; feeling down, depressed or hopeless) form the depression subscale. Each runs from 0 to 6.
The subscale threshold is the one that matters clinically. A score of 3 or more on either subscale is a positive screen for that domain and should prompt administration of the full GAD-7 or PHQ-9. Kroenke and colleagues describe a total of 6 or more as an indicator for further inquiry rather than a diagnostic threshold, and the severity bands below should be read the same way - as descriptive conventions anchored to population percentiles, not as validated diagnostic cutoffs. In the German norming sample a total of 6 sat at the 95.7th percentile and a total of 9 at the 99.1st.
The threshold is not universal across settings. In 2,852 preoperative surgical patients, Kerper and colleagues found that sensitivity at the conventional total of 6 fell to 51.5% and recommended lowering the cutoff to 4, where sensitivity was 80.5% and specificity 80.2%.
Score
Severity
Interpretation
0–2
Normal
Minimal symptoms.
3–5
Mild
Mild distress. An anxiety (items 1-2) or depression (items 3-4) subscale of 3 or more is a positive screen.
6–8
Moderate
Moderate distress.
9–12
Severe
Severe distress. Further assessment recommended.
04 - Validation evidence
How well it performs.
The original 2009 validation was construct-based rather than diagnostic: Kroenke and colleagues did not report sensitivity and specificity against a structured diagnostic interview, testing the scale instead against functional status, disability days and healthcare use. Correlations with the SF-20 were strongest for mental health (r = 0.80), then social functioning (0.52), general health perceptions (0.48), role functioning (0.37), bodily pain (0.36) and physical functioning (0.36) - the gradient you would expect from a valid measure of psychological distress.
Diagnostic accuracy evidence arrived later, and sits at subscale level. In the Greek general population, Christodoulaki and colleagues (2022) found the PHQ-2 at a threshold of 3 detected major depressive disorder with sensitivity 0.77 and specificity 0.94, and the GAD-2 at 3 detected generalised anxiety disorder with sensitivity 0.77 and specificity 0.82. Dropping to a threshold of 2 raised sensitivity (0.87 for any depressive disorder; 0.82 for any anxiety disorder) at a cost to specificity (0.85 and 0.75). Cano-Vindel and colleagues (2018), using a computerised PHQ-4 with 1,052 Spanish primary care patients, likewise found 3 optimal on both subscales, with high sensitivity (0.90 PHQ-2, 0.88 GAD-2) but modest specificity (0.61 for both) - a reminder that in low-prevalence primary care populations most positives will be false.
The factor structure is the best-replicated property. In the German norming sample, confirmatory factor analysis of the two-factor solution returned CFI 0.984, TLI 0.988 and RMSEA 0.027 (90% CI 0.023-0.032), with factor loadings between 0.73 and 0.87, and held invariant across four age-by-gender subsamples. A one-factor solution also fitted acceptably (RMSEA 0.059), which is what justifies reporting a total score at all.
α = 0.85
INTERNAL CONSISTENCY, TOTAL (KROENKE 2009, N=2,149)
84%
VARIANCE EXPLAINED BY TWO-FACTOR SOLUTION (KROENKE 2009)
0.77 / 0.94
PHQ-2 SENSITIVITY / SPECIFICITY AT ≥3 FOR MDD (CHRISTODOULAKI 2022)
1.76 (SD 2.06)
MEAN TOTAL, GERMAN GENERAL POPULATION (LÖWE 2010, N≈5,030)
05 - How it compares
How it compares to the alternatives.
Instrument
Items
Time
When to reach for it
PHQ-4
4
~2 min
Range 0-12. Combined anxiety and depression screen where time is the binding constraint; universal triage. Positive at ≥3 on either subscale.
Range 0-6. Fastest depression-only first-stage screen; the second half of the PHQ-4. Positive at ≥3 (sensitivity 83%, specificity 92% for major depression).
Range 0-21. Anxiety severity grading and treatment monitoring; the natural follow-on from a positive anxiety subscale. ≥10 (sensitivity 89%, specificity 82% for GAD).
Range 0-27. Depression severity, diagnostic support and outcome tracking over time; includes a suicidality item. ≥10 moderate; has published responsiveness data.
Range 0-24. Population-level non-specific distress and epidemiological surveillance rather than individual case-finding. ≥13 for serious mental illness (sensitivity 0.36, specificity 0.96).
Range 0-100. Positively worded wellbeing screen where symptom-framed questions would be poorly received. ≤50 warrants further testing.
06 - When to use it
Right tool, wrong tool.
Reach for it when
-You want to screen everyone rather than only those you already suspect - the two-minute burden makes universal administration realistic.
-Intake, triage or waiting-room self-completion, where the question is which longer instrument to reach for next.
-You need anxiety and depression covered in one form without administering two.
-Repeated brief check-ins in a service where a longer battery would not be completed.
-Physical health settings - the scale has been evaluated in oncology, chronic pain, preoperative and gastrointestinal populations.
Reach for something else when
-You need to measure change. There is no published minimal important difference, and the scale's floor effects work against it. Use the PHQ-9 or GAD-7 instead.
-You need a severity grade to guide treatment intensity. Four items cannot support that; the parent instruments can.
-You need to assess suicide risk. The PHQ-4 contains no suicidality item - unlike the PHQ-9 - so a low total is not reassurance. Use a dedicated instrument such as the ASQ or C-SSRS screener.
-The patient is under 18. Validation runs from age 18 upward; the PHQ-A is a separate adapted instrument for ages 10 to 18.
-You are working preoperatively and intend to use the conventional cutoff of 6 - sensitivity there is roughly 50%.
-You want a diagnosis. A positive PHQ-4 is a prompt to ask more questions, nothing more.
07 - Confidence & precision
Reading the score with care.
Reliability is good for a four-item scale. A pooled analysis of 26 studies (N = 93,466) reported McDonald's omega of 0.85 for the total score, 0.77 for the PHQ-2 and 0.78 for the GAD-2, consistent with the alpha of 0.85 in the original primary care sample.
Combining that reliability estimate with the general-population standard deviation of 2.06 gives a standard error of measurement of approximately 0.8 points, and a reliable change index - the difference required before a change is unlikely to be measurement noise - of roughly 3 points. Both figures are derived from published inputs rather than reported directly in any paper, and should be treated as rough guides only.
No minimal clinically important difference has been established for the PHQ-4. This is a documented gap rather than an oversight in the literature: the 2023 systematic review of the scale's psychometric properties does not list responsiveness among its established properties, and at least one validation study states explicitly that it could not examine responsiveness or smallest detectable change. The practical implication is direct - the PHQ-4 is a screener, not an outcome measure. For tracking treatment response, move to the PHQ-9, which has published responsiveness data.
08 - Limitations
What it cannot tell you.
Modest specificity in primary care. At the recommended subscale threshold of 3, specificity has been reported as low as 0.61. In a population where the base rate of the disorder is low, this means the majority of positive screens will not have the condition. The PHQ-4 identifies who to ask more about; it does not identify who is unwell.
No suicidality item. The PHQ-2 half omits the PHQ-9's item 9. A patient can score 0 and still be at risk.
Floor effects. Most people in the general population score 0 or 1 - 47.6% of the German norming sample scored 1 or below - which compresses the scale at the bottom and limits its ability to detect improvement.
The total obscures the pattern. A total of 6 could be entirely anxiety, entirely depression, or evenly split, with quite different clinical implications. Always report both subscales.
Setting-dependent thresholds. The conventional cutoffs performed poorly in preoperative surgical patients. Local validation is worthwhile before adopting a threshold in an unusual population.
Two weeks only. Chronic, fluctuating or episodic presentations may fall outside the recall window.
Not validated below age 18.
Self-report throughout, with the usual vulnerability to under-reporting where the setting does not feel private.
[5]Kerper LF, Spies CD, Tillinger J, Wegscheider K, Salz AL, Weiss-Gerlach E, Neumann T, Krampe H Screening for depression, anxiety and general psychological distress in preoperative surgical patients: a psychometric analysis of the Patient Health Questionnaire 4 (PHQ-4) (2014) ↩
See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.
We value your privacy
We use cookies to analyse site usage and improve your experience. Analytics and embedded media (e.g. YouTube) only load if you accept. Read our cookie policy.