Skip to main content
    Depression · 8 items · 0–24 · Kroenke K, Strine TW, Spitzer RL, et al. (2009). J Affect Disord.

    Patient Health Questionnaire-8: Scoring, Cutoffs & Interpretation

    Eight-item depression severity measure - the PHQ-9 without the self-harm item.

    PHQ-80 / 8

    Over the last two weeks, how often have you been bothered by…

    Scored locally - nothing leaves this page

    01Little interest or pleasure in doing things.
    02Feeling down, depressed, or hopeless.
    03Trouble falling or staying asleep, or sleeping too much.
    04Feeling tired or having little energy.
    05Poor appetite or overeating.
    06Feeling bad about yourself, or that you are a failure or have let yourself or your family down.
    07Trouble concentrating on things, such as reading the newspaper or watching television.
    08Moving or speaking so slowly that other people could have noticed; or being so fidgety or restless that you have been moving around a lot more than usual.
    0 of 80 / 24

    02 - The clinician's brief

    Last reviewed: · Reviewed by Lotte Kjær Svalberg

    01 - What it measures

    What this scale measures.

    The PHQ-8 is an eight-item self-report measure of current depression covering the two weeks before completion. Its items correspond to eight of the nine DSM criteria for major depressive episode: low mood, anhedonia, sleep disturbance, fatigue, appetite change, feelings of failure or guilt, concentration difficulty, and psychomotor change. Each item is rated 0 (not at all) to 3 (nearly every day), giving a total from 0 to 24.

    It is identical to the PHQ-9 except that the ninth item - thoughts of death or self-harm - is omitted. That omission is the whole point of the instrument: it makes the PHQ-8 usable in postal surveys, telephone interviews and research settings where nobody is on hand to respond immediately to an endorsement of suicidal thinking. Because the ninth item is infrequently endorsed and contributes little to the total score, PHQ-8 and PHQ-9 totals are nearly identical in practice.

    02 - Origin & purpose

    Where it comes from.

    The PHQ-8 was described by Kurt Kroenke and colleagues in 2009, derived directly from the PHQ-9 (Kroenke, Spitzer & Williams, 2001), itself the self-report depression module of the Patient Health Questionnaire developed from the PRIME-MD system. The 2009 paper evaluated the PHQ-8 in the 2006 Behavioral Risk Factor Surveillance System, a random-digit-dialled telephone survey of 198,678 US adults - one of the largest samples any depression measure has been tested in.

    The purpose was epidemiological: a depression measure safe for population surveillance, where a suicide item cannot responsibly be asked without the capacity to act on the answer. Kroenke's team showed that current depression defined by the PHQ-8 diagnostic algorithm and by a simple cutpoint of 10 or more gave near-identical prevalence estimates (9.1% and 8.6% respectively), so either approach can be used.

    03 - Scoring & cutoffs

    How scoring works.

    Sum the eight items for a total of 0-24. The severity bands mirror the PHQ-9: 0-4 none/minimal, 5-9 mild, 10-14 moderate, 15-19 moderately severe, 20-24 severe. A score of 10 or more is the standard threshold for clinically significant depression. Because the omitted ninth item rarely adds points, PHQ-9 cutpoints transfer to the PHQ-8 essentially unchanged, with a small loss of sensitivity (see validation evidence).

    Score
    Severity
    Interpretation
    0–4
    None to minimal
    Minimal or no depressive symptoms.
    5–9
    Mild
    Mild depressive symptoms. Watchful waiting; repeat in two to four weeks.
    10–14
    Moderate
    Moderate depression. Consider a treatment plan.
    15–19
    Moderately severe
    Moderately severe depression. Active treatment usually warranted.
    20–24
    Severe
    Severe depression. Initiate treatment and consider referral.

    04 - Validation evidence

    How well it performs.

    The strongest evidence is an individual-participant-data meta-analysis by Wu and colleagues (2020, Psychological Medicine) covering 54 primary studies, which compared the PHQ-8 and PHQ-9 head-to-head against diagnostic interviews. At the standard cutoff of 10, PHQ-8 sensitivity was lower than PHQ-9 sensitivity by 0.03, and across all cutoffs the difference ranged from 0.00 to 0.05; specificity differed by at most 0.01 at every cutoff. In other words, the two instruments are effectively interchangeable for screening. Internal consistency is high: Cronbach's alpha 0.89 in a psychiatric outpatient validation (Shin et al., 2019).

    The same outpatient study is a useful caution: against a structured diagnostic interview in a psychiatric setting, a cutoff of 10 had sensitivity of only 58.3% with specificity 83.1% - screening cutpoints validated in primary care and population samples perform differently where the base rate and symptom profile differ.

    0.89
    CRONBACH'S α

    Shin 2019, psychiatric outpatients

    -0.03
    SENSITIVITY VS PHQ-9 (≥10)

    Wu 2020, IPD meta-analysis of 54 studies

    ±0.01
    SPECIFICITY VS PHQ-9

    Within 0.01 of the PHQ-9 at all cutoffs (Wu 2020)

    198,678
    VALIDATION SAMPLE

    Adults, 2006 BRFSS (Kroenke 2009)

    05 - How it compares

    How it compares to the alternatives.

    Instrument
    Items
    Time
    When to reach for it
    PHQ-8
    8
    ~2 min
    Population and remote screening where no clinician can respond to a suicide item. Scores track the PHQ-9 closely.
    9
    ~3 min
    The parent measure; adds the self-harm item. Reach for it in clinical settings where suicide-related responses can be followed up the same day.
    2
    <1 min
    Ultra-brief first-stage gate; positives proceed to the PHQ-8/9.
    4
    ~1 min
    Combined depression and anxiety screen when both domains need one instrument.
    10
    ~5 min
    ICD-10-aligned alternative that can produce a diagnostic category as well as a severity score.
    20
    ~5-8 min
    Older epidemiological measure; broader affective content, more reverse-scored items.
    21
    ~5-10 min
    Licensed severity measure for monitoring in specialist care.

    06 - When to use it

    Right tool, wrong tool.

    The PHQ-8 exists for the settings where the PHQ-9's ninth item cannot be asked responsibly. Where a suicide-related response can be acted on the same day, the PHQ-9 remains the better instrument.

    Reach for it when

    • -Population surveys and research datasets where no clinician can respond to a suicide item.
    • -Postal, online or app-based screening pipelines with no same-day follow-up.
    • -Monitoring depressive symptoms where suicide risk is assessed separately by a clinician.
    • -Settings wanting PHQ-9-comparable scores without the ninth item.

    Reach for something else when

    • -Any setting where depression screening is also serving as de facto suicide-risk detection - use the PHQ-9 with a response pathway, or a dedicated instrument such as the ASQ or C-SSRS Screener.
    • -Diagnosis of major depression, which requires a clinical interview.
    • -Children and adolescents - use age-appropriate instruments.
    • -Claiming equivalence in specialist psychiatric populations without local verification (see Shin 2019).

    07 - Confidence & precision

    Reading the score with care.

    No PHQ-8-specific minimal clinically important difference has been established. On the PHQ-9, a change of about 5 points is commonly treated as clinically meaningful (Löwe et al., 2004), and given the near-identity of the two totals this is a reasonable working figure for the PHQ-8 - but state the borrowing explicitly when using it. As with the PHQ-9, single-point movements within a band should not be over-interpreted; the instrument is a screener and severity guide, not a precision measurement of change.

    08 - Limitations

    What it cannot tell you.

    It contains no suicide or self-harm item, so a reassuring PHQ-8 says nothing about suicide risk - this must be assessed separately in clinical use. It is a self-report screener, not a diagnostic instrument. Somatic items (sleep, fatigue, appetite) inflate scores in physical illness.

    Most cutpoint evidence is inherited from the PHQ-9 literature rather than generated for the PHQ-8 itself. Sensitivity at the standard cutoff drops substantially in psychiatric outpatient populations (58.3% in Shin 2019), so it should not be relied on to rule out depression in specialist settings.

    Common questions from clinicians

    FAQ

    References

    Source literature.

    1. [1]Kroenke K, Strine TW, Spitzer RL, Williams JBW, Berry JT, Mokdad AH The PHQ-8 as a measure of current depression in the general population. J Affect Disord. 2009;114(1-3):163-173 (2009)
    2. [2]Wu Y, Levis B, Riehm KE, et al. Equivalency of the diagnostic accuracy of the PHQ-8 and PHQ-9: a systematic review and individual participant data meta-analysis. Psychol Med. 2020;50(8):1368-1380 (2020)
    3. [3]Kroenke K, Spitzer RL, Williams JBW The PHQ-9: validity of a brief depression severity measure. J Gen Intern Med. 2001;16(9):606-613 (2001)
    4. [4]Shin C, Lee S-H, Han K-M, Yoon H-K, Han C Comparison of the usefulness of the PHQ-8 and PHQ-9 for screening for major depressive disorder: analysis of psychiatric outpatient data. Psychiatry Investig. 2019;16(4):300-305 (2019)
    5. [5]Löwe B, Unützer J, Callahan CM, Perkins AJ, Kroenke K Monitoring depression treatment outcomes with the Patient Health Questionnaire-9. Med Care. 2004;42(12):1194-1201 (2004)

    Scale without compromise

    See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.