The Global Assessment of Functioning (GAF) is a single clinician rating, from 0 to 100, of a patient's overall psychological, social and occupational functioning. The scale is divided into ten ten-point bands, each describing a level of symptoms or functioning - from superior functioning in a wide range of activities at the top, down to persistent danger of severely hurting self or others at the bottom. The clinician assigns the value matching either symptom severity or functional impairment, whichever is worse.
That final rule is the GAF's defining quirk: it deliberately mixes two things - how symptomatic someone is and how well they function - into one number. A patient with severe obsessions who holds down a demanding job, and a patient with mild symptoms who cannot leave the house, may land on the same score for different reasons. That compression is what made the GAF fast and universal, and also what eventually got it removed from the DSM.
02 - Origin & purpose
Where it comes from.
The GAF descends from Luborsky's Health-Sickness Rating Scale (1962) via the Global Assessment Scale (GAS), published by Endicott, Spitzer, Fleiss and Cohen in 1976 as a 1–100 rating of overall severity of psychiatric disturbance. A modified version entered DSM-III-R in 1987 as the GAF and became Axis V of the DSM-IV multiaxial system, which made it, for two decades, one of the most-recorded numbers in world psychiatry - used for treatment planning, service eligibility, disability determinations and outcome monitoring.
DSM-5 (2013) abolished the multiaxial system and dropped the GAF, citing its conceptual ambiguity (mixing symptoms, suicide risk and disability in the same descriptors) and questionable psychometrics in routine practice, suggesting the WHO's WHODAS 2.0 as an alternative. The GAF nonetheless survives in many services and registries - including Scandinavian public psychiatry - because nothing equally fast has fully replaced the shared shorthand it provided.
03 - Scoring & cutoffs
How scoring works.
The clinician selects the ten-point band that best matches the patient's current state - rating symptoms or functioning, whichever is worse - then picks a specific value within the band (0 means inadequate information). Broad orientation: scores above 70 indicate at most mild, transient difficulties; 51–70 mild to moderate symptoms or difficulties; 31–50 serious symptoms or serious impairment (a common neighbourhood for patients in secondary care); below 31, major impairment in multiple domains through to danger to self or others. Some services rate symptom-GAF and function-GAF separately to undo the built-in conflation. There are no diagnostic cutoffs, and historical rules of thumb tying particular diagnoses to particular scores (for example, chronic psychosis "belonging" below 40) are exactly that - habits, not psychometrics, and stigmatising ones.
Score
Severity
Interpretation
1–10
Persistent danger or severe impairment
-
11–20
Very severe impairment
-
21–30
Severe impairment
-
31–40
Major impairment
-
41–50
Serious difficulties
-
51–60
Moderate difficulties
-
61–70
Mild difficulties
-
71–80
Transient difficulties
-
81–90
Good functioning
-
91–100
Superior functioning
-
04 - Validation evidence
How well it performs.
The GAF's evidence is genuinely mixed, and honest reporting matters here. Under research conditions, with training and joint calibration, reliability is respectable: studies of trained raters report intraclass correlations around 0.81, and the underlying GAS showed good reliability in its original studies. In routine clinical practice, agreement falls substantially - the 2005 Psychiatric Services special section and Aas's 2010 review both conclude that untrained, time-pressed raters produce ratings too noisy to compare across clinicians or services. Concurrent validity is moderate: GAF correlates with symptom measures and with the GAS lineage it came from, but its single score cannot separate the symptom and function constructs it contains.
ICC ≈ 0.81
Inter-rater reliability achievable with trained raters (outpatient study)
0–100
Score range across ten ten-point bands; higher is better functioning
1987
Entered the DSM as Axis V (DSM-III-R), after the 1976 Global Assessment Scale
2013
Dropped from DSM-5 over conceptual ambiguity and unreliable routine use
05 - How it compares
How it compares to the alternatives.
Instrument
Items
Time
When to reach for it
GAF
1 rating
2–5 min
Fast global functioning shorthand where a service already shares calibration habits.
Brief health status for health-economic evaluation.
06 - When to use it
Right tool, wrong tool.
Reach for it when
-Quick global functioning shorthand within a team that rates together and calibrates regularly
-Longitudinal tracking of the same patient by the same rater
-Contexts that contractually or legally still require a GAF (some registries, disability frameworks)
Reach for something else when
-Comparing scores across clinicians or services without joint training - reliability does not support it
-Separating symptom severity from functional impairment - the score conflates them by design
-New outcome-measurement programmes - WHODAS 2.0 or a symptom scale plus a function measure is better founded
-Inferring diagnosis or risk from a number - anchors mention risk, but the GAF is not a risk assessment
07 - Confidence & precision
Reading the score with care.
No agreed reliable-change threshold exists for the GAF; with routine-practice reliability, differences of a few points between raters are uninterpretable. Within a single trained rater, movement across a ten-point band boundary is a more defensible signal than any within-band change. If decisions hang on the number - eligibility, disability, discharge - corroborate with a domain-specific measure.
08 - Limitations
What it cannot tell you.
The core problems are structural. One number carries symptoms, functioning and risk, so the same score means different things in different patients. Inter-rater reliability in routine practice is poor, and the scale rewards local rating cultures that drift apart. The anchors reflect 1980s assumptions, some of them stigmatising, about what functioning is compatible with which illness. DSM-5 removed it for exactly these reasons. It also captures a single time point with no defined window, and says nothing about direction of change. None of this makes the GAF useless - it makes it a shorthand, whose value depends entirely on shared calibration.
See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.
We value your privacy
We use cookies to analyse site usage and improve your experience. Analytics and embedded media (e.g. YouTube) only load if you accept. Read our cookie policy.