Global & functioning · 1 item · 1–7 · Guy W (1976). ECDEU Assessment Manual, NIMH.
Clinical Global Impression – Severity (CGI-S): Scoring, Cutoffs & Interpretation
Single clinician rating of overall illness severity, anchored to experience with the same diagnosis.
02 - The clinician's brief
Last reviewed:
01 - What it measures
What this scale measures.
The Clinical Global Impression – Severity scale (CGI-S) is a single-item, clinician-rated measure of how mentally ill a patient is right now, on a seven-point scale from 1 (normal, not at all ill) to 7 (among the most extremely ill patients). The rating is made relative to the clinician's accumulated experience with patients who have the same diagnosis, taking into account everything known: history, psychosocial circumstances, symptoms, behaviour, and the impact of symptoms on the ability to function.
Its companion item, the CGI-I (Improvement), rates change since the start of treatment on a parallel seven-point scale from 1 (very much improved) to 7 (very much worse). Together they compress the whole clinical picture into two numbers - which is both the point and the criticism. The CGI is transdiagnostic: the same scale works in psychosis, mood disorders, anxiety and beyond, which makes it one of the few instruments that lets a service compare severity across diagnostic groups.
02 - Origin & purpose
Where it comes from.
The CGI was published in 1976 in the ECDEU Assessment Manual for Psychopharmacology, edited by William Guy for the US National Institute of Mental Health. It was designed for psychopharmacology trials as a standard way for experienced clinicians to record global severity and global change alongside symptom-specific scales, and it remains one of the most widely used outcome measures in psychiatric research - most drug trials still report a CGI outcome.
Its migration into routine practice came later. Busner and Targum's much-cited 2007 paper argued that the CGI is well suited to everyday clinical work precisely because it demands no kit, no licence and under a minute of time, while forcing the clinician to commit to an explicit global judgement that can be tracked visit to visit. Many services now use the CGI-S as a lightweight, cross-diagnostic severity register - including for sharing a common language about patients across teams.
03 - Scoring & cutoffs
How scoring works.
There is nothing to sum: the clinician assigns one value from 1 to 7 at the time of assessment (0 is recorded when severity was not assessed). The anchor question is "how ill is this patient compared with all the patients with this diagnosis I have known?" - not "compared with the general population". Linking studies give the numbers clinical meaning: in schizophrenia research, CGI-S 3 (mildly ill) corresponds to a PANSS total of roughly 58, and CGI-S 5 (markedly ill) to roughly 88; a fall of about 25% in PANSS total corresponds to a one-step improvement on the CGI-S. There are no diagnostic cutoffs - the CGI-S is a severity and tracking measure, not a screen.
Score
Severity
Interpretation
1–1
Normal, not at all ill
-
2–2
Borderline mentally ill
-
3–3
Mildly ill
-
4–4
Moderately ill
-
5–5
Markedly ill
-
6–6
Severely ill
-
7–7
Among the most extremely ill patients
-
04 - Validation evidence
How well it performs.
Because it is a single global judgement, classical internal-consistency statistics do not apply; the evidence base concerns convergent validity and prognostic value. Leucht and colleagues' linking studies found CGI-S correlated with PANSS totals at r = 0.61 at baseline, rising to about 0.73 after eight weeks of treatment, and established the PANSS/BPRS equivalents above. Equivalent linking work exists for depression and other disorders. A 2023 real-world registry analysis found CGI-S predicted subsequent psychiatric hospitalisation across diagnoses. Reliability is the weak flank: agreement between raters depends on shared calibration and experience, and recent methodological reviews (e.g. Worden, 2024) urge more careful use where raters are not trained together.
r = 0.61–0.73
Correlation between CGI-S and PANSS total across treatment weeks in linking studies (Leucht et al.)
~25%
Reduction in PANSS total corresponding to a one-step CGI-S improvement (Leucht et al., 2006)
<1 min
Time to administer; no materials or licence required
1976
Published in the NIMH ECDEU manual; public domain ever since
05 - How it compares
How it compares to the alternatives.
Instrument
Items
Time
When to reach for it
CGI-S
1 item
<1 min
Cross-diagnostic global severity; tracking and team communication when a shared, quick measure is needed.
Classic clinician-rated depression measure for structured follow-up.
06 - When to use it
Right tool, wrong tool.
Reach for it when
-Tracking global severity across visits in any diagnosis, with near-zero administration burden
-Comparing severity across diagnostic groups within a service or register
-Anchoring symptom-scale scores to a clinically meaningful global judgement
-Trials and audits needing a universal outcome measure alongside disorder-specific scales
Reach for something else when
-Diagnosis or screening - the CGI-S assumes the diagnostic assessment has already happened
-Fine-grained symptom profiling - a single number cannot say what changed
-Comparisons across raters who have not calibrated with each other
-Ratings by clinicians without experience of the relevant patient group - the scale's anchor IS that experience
07 - Confidence & precision
Reading the score with care.
The CGI-S has no standard error of measurement in the usual sense; its precision is bounded by rater calibration. Practical guidance from the linking literature: treat a one-step change as clinically meaningful when the same clinician rates the same patient over time, and be sceptical of one-step differences between different raters. Where precision matters - research, medicolegal work - pair the CGI-S with a disorder-specific scale and report both.
08 - Limitations
What it cannot tell you.
Everything about the CGI-S follows from its brevity. It is subjective, and anchored to each clinician's private case series, so systematic differences between raters are built in. It collapses symptoms, function and risk into one number without saying which moved. It offers no guidance for the rating beyond the anchor labels, so untrained raters drift. And because it assumes familiarity with the diagnostic group, it performs worst exactly where a junior clinician might want it most. Recent critiques argue it is too often used as a stand-alone outcome measure in research settings where its psychometric limits are underappreciated.
See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.
We value your privacy
We use cookies to analyse site usage and improve your experience. Analytics and embedded media (e.g. YouTube) only load if you accept. Read our cookie policy.