Wellbeing · 0 items · 0–0 · Developed by John E. Ware Jr and colleagues for the RAND Medical Outcomes Study (1992). Licensed by QualityMetric (IQVIA).
Licensed instrument
SF-36: scoring, cutoffs & interpretation
A licensed 36-item measure of health-related quality of life across eight domains, with physical and mental component summaries. Information resource only - items are not hosted here.
The SF-36 is a generic measure of health-related quality of life: how a person's health affects their everyday functioning and wellbeing, regardless of diagnosis. Its 36 items produce eight scale scores, each on a 0-100 scale where higher means better health: physical functioning, role limitations due to physical health, bodily pain, general health perceptions, vitality, social functioning, role limitations due to emotional problems and mental health. A single additional item asks about change in health over the past year and is reported separately.
Because it is generic, the SF-36 can compare the burden of very different conditions (for example depression against heart failure) and track the broad impact of treatment on function rather than on symptoms alone. The eight scales can be aggregated into two norm-based summary scores, the Physical Component Summary (PCS) and the Mental Component Summary (MCS), scaled so that the general population has a mean of 50 and standard deviation of 10. In psychiatric settings the mental health scale and MCS are the most relevant, and the MCS has been used as a depression screen.
02 - Origin & purpose
Where it comes from.
The SF-36 was constructed for the Medical Outcomes Study (MOS), a large RAND-led study of how physician practice style and health system characteristics affected patient outcomes in the United States. John Ware and Cathy Sherbourne published the conceptual framework and item selection in Medical Care in 1992, with companion papers by McHorney and colleagues in 1993 and 1994 establishing validity, data quality and reliability across 24 patient subgroups. The 36 items were chosen from the longer MOS instruments to cover eight health concepts in a form that could be self-administered in five to ten minutes.
Two lineages emerged from the same items. The RAND Corporation released the original items as the RAND 36-Item Health Survey 1.0 with a simple averaging method of scoring, free of charge. Ware and colleagues developed the SF-36 through the Medical Outcomes Trust and later QualityMetric, adding norm-based scoring, the PCS and MCS summaries, and in 1996 a second version (SF-36v2) with five-level response options for the role items and simplified wording. The SF-12 (1996) reproduces the PCS and MCS from a 12-item subset. The SF-36 has been translated into dozens of languages through the IQOLA project and is one of the most widely cited outcome measures in medicine.
03 - Scoring & cutoffs
How scoring works.
Each SF-36 item is recoded so that a higher value means better health, the items in each scale are summed, and the raw sum is transformed linearly to a 0-100 range. In version 1 two items are recoded non-linearly (the general health rating and the pain items), which is one of the differences from RAND-36 scoring. A scale is scored if the respondent answered at least half its items; missing items are replaced with the person's mean on the answered items in that scale. The PCS and MCS are then calculated by weighting the eight scale z-scores with factor coefficients from the US general population and rescaling to a mean of 50 and SD of 10; in SF-36v2 the eight scales are also reported as norm-based T-scores, so a score of 40 on any scale is one standard deviation below the population mean.
There are no diagnostic cutoffs on the eight 0-100 scales; they are interpreted against age- and sex-matched population norms or the patient's own baseline. Two screening thresholds are used for mental health: a Mental Health scale score of 52 or below, and an MCS of 42 or below, both proposed in the 1994 summary-scales manual as indicators of possible depressive disorder. The standard form uses a four-week recall period; an acute one-week version exists for short-interval studies. Scoring requires licensed software or the published algorithms, and PCS/MCS should not be estimated by simply averaging the scales.
04 - Validation evidence
How well it performs.
In the MOS sample of 3,445 patients, McHorney and colleagues found internal consistency reliability exceeding 0.80 for all scales except social functioning, with Cronbach's alpha ranging from 0.65 to 0.94 across scales and subgroups; the two-item social functioning scale was the least reliable. Floor effects were negligible except on the role scales, and ceiling effects were marked on role-physical, role-emotional and social functioning in version 1, which motivated the five-level role items of SF-36v2. Two-week test-retest reliability was demonstrated in UK primary care by Brazier and colleagues in 1992. The SF-12 reproduces the SF-36 PCS and MCS with R-squared of 0.91 and 0.92 in the US general population. As a depression screen in older women, an MCS of 42 or below had sensitivity 71% and specificity 82% against DSM-III-R diagnosis; the MH <=52 threshold was more specific (92%) but less sensitive (58%), and neither performed well for anxiety disorders. In 532 older adults with major depression, the mental health scale and MCS were the SF-36 components most responsive to change in depression severity over six weeks.
alpha 0.65-0.94
Internal consistency
Across eight scales and 24 patient subgroups; all scales except social functioning exceeded 0.80 - McHorney et al. 1994
Sens 71% / Spec 82%
Depression screen (MCS <=42)
Against DSM-III-R depression in 586 women aged 70-84 - Silveira et al. 2005
Sens 58% / Spec 92%
Depression screen (MH <=52)
Same sample; more specific, less sensitive - Silveira et al. 2005
R-squared 0.91 / 0.92
SF-12 reproduction of PCS/MCS
US general population, n = 2,333 - Ware, Kosinski & Keller 1996
05 - How it compares
How it compares to the alternatives.
Instrument
Items
Time
When to reach for it
SF-36
36
5-10 min
Eight-scale profile plus PCS/MCS with extensive international norms; licensed - not hosted here.
Free, very brief positive wellbeing index for routine monitoring.
06 - When to use it
Right tool, wrong tool.
Reach for it when
-Measuring the broad functional and wellbeing impact of illness or treatment, alongside symptom scales.
-Comparing burden across diagnoses or against population norms.
-Clinical trials and service evaluations where PCS/MCS and international norms are required and a licence is in place.
-Screening for probable depression in medical populations using MCS <=42 when a dedicated depression scale is not available.
Reach for something else when
-Routine psychiatric outcome monitoring without a licence - use the free RAND-36, WHODAS 2.0 or WHO-5.
-Diagnosing depression or anxiety - the MCS is a coarse screen; use the PHQ-9, MDI or GAD-7.
-Health economic evaluation needing a single utility index - use the EQ-5D-5L or map to SF-6D.
-Children under 14 or people with severe cognitive impairment, for whom it is not validated.
07 - Confidence & precision
Reading the score with care.
Reliability differs substantially across the eight scales, so precision does too. With reliabilities around 0.90 and population SDs of roughly 20-25 points, the SEM on the better scales is about 7 points, but on the two-item social functioning scale (alpha ~0.65-0.75) it can exceed 12 points. For the PCS and MCS, the widely used rule of thumb from Norman and colleagues is that a minimally important difference is about half a standard deviation, roughly 5 points; at the individual level, changes of about 6.5 (PCS) and 7.9 (MCS) points have been proposed as reliably beyond error, whereas for group comparisons differences of 2-3 points can be meaningful. On the 0-100 scales, anchor-based studies commonly place the minimal important difference between 3 and 5 points. Individual changes smaller than these should not be over-interpreted.
08 - Limitations
What it cannot tell you.
The SF-36 is licensed and cannot be reproduced or hosted without agreement from QualityMetric, and scoring the PCS and MCS requires licensed software or the published algorithms. It is not preference-based, so it does not yield a single utility value for economic evaluation without mapping to the separately licensed SF-6D. Version 1 has ceiling effects on the role scales and the two-item social functioning and pain scales are imprecise. Norms are country-specific and the factor structure underlying PCS and MCS does not hold equally in all cultures; the PCS/MCS weighting can also produce counter-intuitive results, for example a physical illness lowering MCS. It is designed for people aged 14 and over and is not validated in children. As a generic measure it is less sensitive to change than disease-specific scales and does not replace symptom measures in psychiatry.
09 - Licensing, explained
How licensing works.
The SF-36 and SF-36v2 are copyrighted and trademarked (Medical Outcomes Trust, Health Assessment Lab, QualityMetric and Optum) and licensed through QualityMetric, now part of IQVIA. Commercial use requires a paid licence; unfunded academic research can register for royalty-free use through QualityMetric's Office of Grants and Scholarly Research. Because of this, Aisel does not host the SF-36 items. The RAND 36-Item Health Survey 1.0 (RAND-36) uses the same original Medical Outcomes Study items, is free to use with a RAND credit line, and is available in the Aisel scale library.
See how Aisel removes friction where it costs most. A 20-minute walkthrough tailored to your clinic.
We value your privacy
We use cookies to analyse site usage and improve your experience. Analytics and embedded media (e.g. YouTube) only load if you accept. Read our cookie policy.